DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
Researchers have released DialToM, a new Theory of Mind benchmark that tests large language models’ ability to predict dialogue trajectories based solely on mental‑state profiles, without any conversational context. The study finds that while models can infer mental states (Literal ToM) with high accuracy, they struggle to use those inferences for social forecasting (Functional ToM), a gap that a human expert closes entirely. The benchmark, detailed in arXiv:2604.20443v3, highlights a clear disparity between human and AI performance in state‑driven dialogue prediction.

Researchers have released DialToM, a new Theory of Mind benchmark that tests large language models’ ability to predict dialogue trajectories based solely on mental‑state profiles, without any conversational context. The study finds that while models can infer mental states (Literal ToM) with high accuracy, they struggle to use those inferences for social forecasting (Functional ToM), a gap that a human expert closes entirely. The benchmark, detailed in arXiv:2604.20443v3, highlights a clear disparity between human and AI performance in state‑driven dialogue prediction.
Sources
- arXiv cs.LG — DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories
由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。