Researchers have identified a bias in Video Large Language Models (Video-LLMs) where they overly focus on a single "anchor frame" during generation, leading to hallucinations. To mitigate this issue, they propose Decoder-side Temporal Rebalancing (DTR), a training-free method that encourages models to consider more frames evenly, improving robustness against hallucinations without compromising video understanding performance or efficiency.
Read the full article at arXiv cs.CV (Vision)
Want to create content about this topic? Use Nemati AI tools to generate articles, social posts, and more.



