Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video
A new study demonstrates that a vision‑language model can estimate the dynamic, triaxial hand forces exerted during manual material handling from ordinary RGB video, eliminating the need for instrumented objects. Using a pipeline that combines task‑specific text prompts, visual region‑of‑interest detection, and transformer‑based temporal regression, researchers tested 35 participants performing lifting, carrying, pushing, and pulling tasks with 6‑12 kg boxes and achieved reliable force predictions across seven camera‑view conditions. The work suggests a scalable, non‑intrusive method for assessing occupational physical exposure and injury risk in real‑world settings.

A new study demonstrates that a vision‑language model can estimate the dynamic, triaxial hand forces exerted during manual material handling from ordinary RGB video, eliminating the need for instrumented objects. Using a pipeline that combines task‑specific text prompts, visual region‑of‑interest detection, and transformer‑based temporal regression, researchers tested 35 participants performing lifting, carrying, pushing, and pulling tasks with 6‑12 kg boxes and achieved reliable force predictions across seven camera‑view conditions. The work suggests a scalable, non‑intrusive method for assessing occupational physical exposure and injury risk in real‑world settings.
Sources
- arXiv cs.AI — Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video
由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。