PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
PathAgentBench is a new benchmark introduced for evaluating evidence-seeking vision-language models (VLMs) on whole-slide pathology images. Unlike existing benchmarks that focus on pre-cropped patches or slide features, PathAgentBench assesses models' ability to interpret and verify evidence directly from gigapixel WSIs across magnifications. The benchmark includes 1,822 TCGA WSIs and 17,135 diagnostic paths annotated by ten board-certified pathologists. It evaluates four key capabilities: image-to-text matching for interpretation, text-to-image retrieval for verification, diagnostic-region localization, and multi-scale reasoning. For more details, refer to the original paper on arXiv:2607.19261v3.

PathAgentBench is a new benchmark introduced for evaluating evidence-seeking vision-language models (VLMs) on whole-slide pathology images. Unlike existing benchmarks that focus on pre-cropped patches or slide features, PathAgentBench assesses models' ability to interpret and verify evidence directly from gigapixel WSIs across magnifications. The benchmark includes 1,822 TCGA WSIs and 17,135 diagnostic paths annotated by ten board-certified pathologists. It evaluates four key capabilities: image-to-text matching for interpretation, text-to-image retrieval for verification, diagnostic-region localization, and multi-scale reasoning. For more details, refer to the original paper on arXiv:2607.19261v3.
Sources
- arXiv cs.AI — PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。