pathology-cot-trains-visual-chain-of-thought-agents-using-expert-whole-slide-diagnoses
Pathology-CoT Trains Visual Chain-of-Thought Agents Using Expert Whole-Slide Diagnoses

Pathology-CoT Trains Visual Chain-of-Thought Agents Using Expert Whole-Slide Diagnoses

Whole-slide imaging (WSI) promises pathology at scale, but today’s diagnostic AI still struggles with the messy, interactive way experts actually examine tissue. A pathologist doesn’t simply “classify” an image; they navigate, change magnification, zoom in on suspicious regions, and iteratively connect visual cues to clinical reasoning. That tacit, experience-driven workflow has been largely missing from training data—leaving many agentic systems unable to move beyond static vision tasks.

In a new study, researchers propose Pathology‑CoT, a framework designed to transform expert viewing behavior into supervision that an AI agent can follow. The central idea is to capture the sequence of expert attention during routine WSI review and convert it into structured instructions, so the model learns not only what to detect, but how to look.

The work introduces an “artificial intelligence session recorder” that passively logs how clinicians move through standard slide viewers. Raw interaction traces—such as navigation patterns and viewing changes—are translated into standardized behavioral commands and bounding-box annotations. This creates a bridge between human viewing trajectories and machine-readable guidance.

Next, the team adds a human-in-the-loop review stage. Draft rationales produced by AI are checked and refined, and each step is paired with supervision that explicitly answers two questions: where the model should focus and why that focus matters. The authors report that this pipeline enables substantially faster labeling—roughly sixfold compared with conventional annotation approaches.

Using the resulting dataset, the team builds Pathology‑o3, a two-stage agent. The first stage proposes candidate regions of interest on the slide, while the second performs behavior-guided reasoning that mirrors an expert’s stepwise exploration. Instead of treating the WSI as a single inference problem, the agent decomposes diagnosis into an interpretable sequence of attention.

On gastrointestinal lymph node metastasis detection, Pathology‑o3 outperformed state-of-the-art vision–language models. Importantly, performance improvements were consistent across multiple vision–language model backbones, suggesting the framework is not tied to a single architecture.

The approach also generalizes beyond development data: the agent maintained strong performance on an independent external validation cohort. Together, these results position Pathology‑CoT as a step toward more reliable and explainable WSI agents that better reflect real expert practice.

By combining unobtrusive behavior logging, curated “where-to-look/why-it-matters” supervision, and agentic reasoning, Pathology‑CoT aims to close a key gap in AI training—bringing the logic of diagnosis back into the way models explore tissue.

Subject of Research: Whole-slide image (WSI) pathology diagnosis; agentic visual reasoning; expert behavior learning.

Article Title: Pathology‑CoT: learning visual chain-of-thought agents from expert whole-slide image diagnosis behaviour.

Article References: Wang, S., Wu, R., Herndon, C. et al. Pathology‑CoT: learning visual chain-of-thought agents from expert whole-slide image diagnosis behaviour. Nat. Biomed. Eng (2026). https://doi.org/10.1038/s41551-026-01739-y

Image Credits: AI Generated

DOI: https://doi.org/10.1038/s41551-026-01739-y

Keywords:

Tags: AI training for interactive pathology workflowsbehavior-based AI training in histopathologycapturing clinician attention in pathologydynamic tissue analysis with AIexpert-driven tissue examination modelinghuman-in-the-loop pathology AI refinementpathology diagnosis automation with expert behaviorpathology whole-slide imaging analysisslide viewer interaction loggingstructured supervision for diagnostic AItranslating visual cues into AI guidancevisual chain-of-thought in medical AI