OneStreamer combines perception and memory to respond when evidence is sufficient.
4BModel parameters
8/8Benchmarks led
1.16MTraining records
A moment passes. The evidence stays.
The reading corner is no longer in view.
7 / 7
User · Response
Where can my kids go to read?
Memory
More examples +
Recalling earlier observationsTimely responses
See. Remember. Know when to respond.
Details observed in a live video may become relevant only after their frames have left the model’s visual context. OneStreamer preserves these observations as time-grounded factual memory. It combines recent visual evidence with proactively generated captions and learns when the available evidence is sufficient to record or respond.
Read the full abstract +
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
How OneStreamer works.
One causal visual–language sequence brings together current frames, past caption records, and the ongoing dialogue.
Turn observations into factual memory.
Proactive Hierarchical Caption Memory records local details with </Observe> and completed events with </Summary>. These records remain available after the source frames leave the recent visual window.
Local details</Observe>
Event summaries</Summary>
Recent visual window + caption history
Learn the moments that call for an output.
Proactive State Transition Learning preserves all output anchors and selects representative state changes and state-persistence positions. All control tokens remain in the sequence; only the state loss is masked.
OneStreamer-1M: Grounded in evidence.
Synthesized streaming captions and QA join cleaned open-source data in a shared format. Every target is aligned with evidence available at its assigned time.
1,160,004streaming annotation records
Perception & Memory QA630,850
Proactive Interaction271,570
Proactive QA197,584
Proactive Caption Memory60,000
Data synthesis
One 4B model. Eight leading results.
OneStreamer achieves the highest aggregate score among the compared methods on every evaluated benchmark.
OneStreamer (4B) Qwen3-VL base (4B)
Online video benchmark results.
Better memory. Well-timed responses.
Ablations isolate the value of retaining caption records and choosing informative state supervision.
63.5 →71.6
Remember beyond the visual window.
OVOBench Backward ASI improves with retained caption memory, using the same checkpoint and Recent-16 window as FIFO.
27.5%
Supervise fewer, better-chosen states.
PSTL outperforms dense state supervision and a supervision-matched random baseline on three proactive-response benchmarks.
Memory and perception (PHCM)+
Effects of Proactive Hierarchical Caption Memory
Variant
Visual context
OVO Backward ASI ↑
OVO Backward EPM ↑
OVO Real-Time ↑
OVO Overall ↑
StreamingBench RT ↑
Full
Full history
67.6
63.0
70.9
70.9
80.5
FIFO
Recent-16
63.5
62.0
80.9
70.1
86.3
PHCM
Recent-16 + captions
71.6
62.6
81.4
72.1
86.9
Response timing (PSTL)+
Ablation of state-token supervision
State selection
Loss
Supervision
ProactiveVQA ↑
OmniMMI ↑
OVO-Timing ↑
All state tokens
CE
100.0%
26.1
30.8
1.5
All state tokens
Focal
100.0%
47.0
32.8
26.5
Random Sparse
CE
27.5%
26.6
27.4
1.4
Transition Only
CE
18.1%
46.7
27.4
15.3
PSTL
CE
27.5%
48.7
36.6
41.6
Answer-stage efficiency
On a 360-second OVOBench sample, PHCM uses 93.1% fewer context tokens than retaining the full visual history.
124mstime to first token at the answer stage
Answer-stage efficiency
Strategy
Context tokens ↓
GPU memory ↓
TTFT ↓
Full
62,094
25.18 GB
4.560 s
FIFO
3,036
9.69 GB
0.094 s
PHCM
4,308
9.98 GB
0.124 s
Single H200. FIFO and PHCM use 16 recent frames; PHCM adds precomputed caption memory.
Online memory construction+
On a 360-second StreamingBench clip, PHCM averages 0.636 seconds per update, below the one-second input interval. This includes video decoding, preprocessing, input construction, vision encoding, and language-model inference.
Online memory construction on a 360-second StreamingBench clip
Strategy
Processing time ↓
Mean update ↓
Peak memory ↓
Peak input tokens ↓
FIFO (forced silence)
170.00 s
0.472 s
10.233 GiB
7,914
Full (forced silence)
1444.75 s
4.013 s
29.754 GiB
87,535
PHCM
228.96 s
0.636 s
10.327 GiB
8,362
Single H200, one update per second. FIFO and PHCM retain 128 frames (32 s at 4 FPS); Full retains all observed frames. FIFO and Full emit silence, while PHCM generates caption memory.
Citation
@misc{zeng2026onestreamerunifyingperceptionmemory,
title={OneStreamer: Unifying Perception, Memory, and Proactive
Response in Streaming Video Interaction},
author={Xiangyu Zeng and Yuandong Yang and Zhiqiu Zhang and
Yuhan Zhu and Xinhao Li and Qingyi Si and Dingyu Yao and
Changlian Ma and Haoran Chen and Xinyu Chen and Yansong Shi and
Junhao Zhou and Yifei Li and Jun Zhang and Chuanyu Qin and
Chenxu Yang and Xinlei Yu and Kun Ouyang and Yuchen Shao and
Qianshan Wei and Changhai Zhou and Jun Gao and Jiaqi Wang and
Limin Wang},
year={2026},
eprint={2610.01762},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.01762}
}