Video-Based Large Language Model for Enhanced Temporal Understanding
Loading...
Files
Date
Contributor
Advisor
Editor
Performer
Department
Instructor
Depositor
Speaker
Researcher
Consultant
Interviewer
Interviewee
Narrator
Transcriber
Annotator
Journal Title
Journal ISSN
Volume Title
Publisher
Journal Name
Volume
Number/Issue
Starting Page
5121
Ending Page
Alternative Title
Abstract
This paper presents a novel video-based large language model (LLM) designed to tackle key challenges in temporal reasoning and multimodal integration for video understanding. Unlike existing approaches, our model introduces Enhanced Temporal Position Embeddings (ETPE) to capture long-range temporal dependencies and a Modal Fusion Bridge (MFB) to dynamically integrate visual and auditory modalities through bidirectional attention. Additionally, we implement an adaptive frame sampling strategy that minimizes redundancy while retaining motion-rich segments to enhance computational efficiency. Our system achieves state-of-the-art performance, recording 92.6% accuracy on Kinetics-400 and 90.4% on Kinetics-600—using only 5% of the training data. These innovations lead to significant gains in efficiency and accuracy across tasks such as video question answering, event localization, and real-time processing. We further demonstrate the practical applicability of our model through a Gradio-based interactive interface, enabling real-world use cases such as sports video analysis and accessibility support. Overall, this work establishes a scalable and efficient framework for multimodal video understanding systems.
Description
Citation
Extent
8 pages
Format
Type
Conference Paper
Geographic Location
Time Period
Related To
Proceedings of the 59th Hawaii International Conference on System Sciences
Related To (URI)
Table of Contents
Rights
Attribution-NonCommercial-NoDerivatives 4.0 International
Rights Holder
Catalog Record
Local Contexts
Collections
Email libraryada-l@lists.hawaii.edu if you need this content in ADA-compliant format.
