TimeLens2

Generalist Video Temporal Grounding with Multimodal LLMs

A unified 4B/8B temporal grounder that predicts variable-cardinality evidence sets across video lengths, domains, query forms, and viewpoints.

47.7Avg. mIoU · 4B
48.0Avg. mIoU · 8B
+13.04B backbone gain
+18.18B backbone gain
93KVerified examples
23,793Curated videos

Overview

One interface for temporal evidence across heterogeneous video settings.

TimeLens2 represents temporal evidence as an interval set throughout supervision and optimization. The same generative interface handles single-span, multi-span, question-form, long-video, and egocentric grounding.

TimeLens2 results across seven temporal grounding benchmarks and representative query types
Figure 1. TimeLens2 unifies seven benchmarks spanning short, long, multi-span, question-form, highlight, and egocentric grounding.

TimeLens2-93K

Turn long-video annotation into a sequence of verifiable decisions.

Caption-derived proposals narrow the search space; independent local grounding, temporal consensus, semantic checks, and boundary refinement determine the final interval sets.

10.2 minaverage video duration
23,793retained videos
93,232grounding instances
12,091multi-span instances
Six-stage TimeLens2-93K data construction pipeline and retained corpus statistics
Figure 2. Six stages separate candidate construction from label determination, ending in verified and locally refined interval sets.Full resolution

Method

Align training with the structure of temporal evidence.

Sparse, variable-cardinality intervals require supervision and objectives that preserve their set-level geometry.

01

Verified supervision

Caption-derived proposals are independently localized, checked by cross-agent consensus, semantically verified, and boundary-refined.

02

Long-context generalist

TimeLens2-93K and contexts up to 100K tokens teach one 4B or 8B model to retain sparse evidence through distractor-heavy videos.

03

Temporal Wasserstein reward

Exact one-dimensional W₁ over merged interval support complements tIoU with dense, matching-free credit for near and distant errors.

RTW = exp(−W₁ / (|merge(Y)| + ε))
Temporal Wasserstein reward distinguishes near misses and avoids interval matching failures
Figure 3. Temporal Wasserstein reward distinguishes near misses even when tIoU is zero and is invariant to redundant interval fragmentation.
75.8%all-zero-tIoU groups receive useful ordering
13.8 → 3.6%constant-reward groups after adding temporal W₁
4.0×higher within-group reward variance

Results

Compact models with broad temporal range.

TimeLens2-4B TimeLens2-8B Qwen3-VL-235B-A22B Qwen3.5-397B-A17B
CharadesShort action
57.7
58.6
47.8
47.5
ActivityNetEvent-level
59.0
58.6
52.2
53.8
QVHighlightsHighlight
69.3
70.2
64.6
65.8
VUE-TRLong / multi
53.2
53.5
43.1
42.2
VUE-TR-V2Long / user
48.1
47.7
34.1
34.5
MomentSeekerQuestion-form
27.9
28.5
23.7
23.3
Ego4D-NLQEgocentric
18.6
19.0
12.7
14.5

Qualitative analysis

The right event, not merely a plausible scene.

The model remains calibrated on long videos, preserves multiple sparse moments, grounds OCR evidence, and rejects visually similar distractors.

Qualitative comparison on long-video, multi-span, question-form, and egocentric temporal grounding
Figure 4. TimeLens2 retrieves late, recurrent, answer-bearing, and first-person evidence while rejecting plausible distractors.Full resolution

Supplementary examples

Seven benchmarks, seven temporal reasoning challenges.

Additional cases isolate state transitions, conjunctive events, role binding, repeated evidence, sparse long-range relations, anomaly semantics, and first-person preconditions.

Additional qualitative TimeLens2 examples across all seven temporal grounding benchmarks
Supplementary Figure S1. Both TimeLens2 scales remain precise and complete across all seven evaluation regimes.Full resolution