Conference item icon

Conference item

Understanding adaptive frame selection in medical video large language models through task-derived evidence supervision

Abstract:
Medical Video Large Language Models (Medical Video LLMs) typically process only a limited number of video frames because of computational constraints, making temporal frame selection a critical compo- nent of efficient medical video understanding. Existing approaches pri- marily rely on fixed sampling strategies or generic frame importance estimation, while overlooking the rich supervisory signals already avail- able in medical video annotations. In this work, we investigate adaptive frame selection through Task-Derived Evidence Supervision (TDES), a framework that converts existing task annotations into frame-level rele- vance supervision for training a lightweight Task-Aware Adaptive Video Sampler (TAVS). The learned sampler operates independently of the downstream Medical Video LLM and can therefore be integrated into existing inference pipelines without architectural modification. We con- duct a controlled comparison of evidence-guided and conventional frame selection strategies on the MedVidBench benchmark using two repre- sentative Medical Video LLM backbones, Qwen2.5-VL-7B-Instruct and UAI-NEXUS-MedVLM-1.0a-7B-RL, and compare them with uniform, heuristic, contiguous-window, and learned sampling baselines, as well as prompt-based temporal metadata augmentation. Our experiments show that evidence-guided frame selection yields the strongest Surgical Assess- ment performance on Qwen while remaining competitive across other tasks. In contrast, MedVLM exhibits remarkable robustness to differ- ent frame selection strategies, and explicit temporal metadata provides only marginal benefit for either backbone. Rather than identifying a universally superior frame selection strategy, our findings reveal that the effectiveness of adaptive temporal sampling depends strongly on the underlying Medical Video LLM backbone. This study suggests that fu- ture progress in temporal sampling should consider both the supervision strategy and the characteristics of the downstream backbone.
Publication status:
Accepted
Peer review status:
Peer reviewed

Actions

Authors

More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author


Publisher:
Springer
Series:
Lecture Notes in Computer Science
Acceptance date:
2026-08-06
Event title:
19th European Conference on Computer Vision (ECCV 2026)
Event location:
Malmö, Sweden
Event website:
https://eccv2024.ecva.net/Conferences/2026
Event start date:
2026-09-08
Event end date:
2026-09-12


Language:
English
Keywords:
Pubs id:
2452093
Local pid:
pubs:2452093
Deposit date:
2026-08-19
ARK identifier:

Terms of use


Views and Downloads






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP