Conference item
Understanding adaptive frame selection in medical video large language models through task-derived evidence supervision
- Abstract:
- Medical Video Large Language Models (Medical Video LLMs) typically process only a limited number of video frames because of computational constraints, making temporal frame selection a critical compo- nent of efficient medical video understanding. Existing approaches pri- marily rely on fixed sampling strategies or generic frame importance estimation, while overlooking the rich supervisory signals already avail- able in medical video annotations. In this work, we investigate adaptive frame selection through Task-Derived Evidence Supervision (TDES), a framework that converts existing task annotations into frame-level rele- vance supervision for training a lightweight Task-Aware Adaptive Video Sampler (TAVS). The learned sampler operates independently of the downstream Medical Video LLM and can therefore be integrated into existing inference pipelines without architectural modification. We con- duct a controlled comparison of evidence-guided and conventional frame selection strategies on the MedVidBench benchmark using two repre- sentative Medical Video LLM backbones, Qwen2.5-VL-7B-Instruct and UAI-NEXUS-MedVLM-1.0a-7B-RL, and compare them with uniform, heuristic, contiguous-window, and learned sampling baselines, as well as prompt-based temporal metadata augmentation. Our experiments show that evidence-guided frame selection yields the strongest Surgical Assess- ment performance on Qwen while remaining competitive across other tasks. In contrast, MedVLM exhibits remarkable robustness to differ- ent frame selection strategies, and explicit temporal metadata provides only marginal benefit for either backbone. Rather than identifying a universally superior frame selection strategy, our findings reveal that the effectiveness of adaptive temporal sampling depends strongly on the underlying Medical Video LLM backbone. This study suggests that fu- ture progress in temporal sampling should consider both the supervision strategy and the characteristics of the downstream backbone.
- Publication status:
- Accepted
- Peer review status:
- Peer reviewed
Actions
Authors
- Publisher:
- Springer
- Series:
- Lecture Notes in Computer Science
- Acceptance date:
- 2026-08-06
- Event title:
- 19th European Conference on Computer Vision (ECCV 2026)
- Event location:
- Malmö, Sweden
- Event website:
- https://eccv2024.ecva.net/Conferences/2026
- Event start date:
- 2026-09-08
- Event end date:
- 2026-09-12
- Language:
-
English
- Keywords:
- Pubs id:
-
2452093
- Local pid:
-
pubs:2452093
- Deposit date:
-
2026-08-19
- ARK identifier:
Terms of use
- Notes:
- This paper will be presented at the 19th European Conference on Computer Vision (ECCV 2026), 8th-12th September 2026, Malmö, Sweden.
If you are the owner of this record, you can report an update to it here: Report update to this record