Instructions to use zeyun-zhong/StreamTTT-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zeyun-zhong/StreamTTT-4B with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForCausalLM processor = AutoProcessor.from_pretrained("zeyun-zhong/StreamTTT-4B", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("zeyun-zhong/StreamTTT-4B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
StreamTTT-4B
Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
📄 Paper · 💻 Code · 📊 Data
Streaming video assistants are pulled in two directions. Keeping only a short window of recent frames gives sharp perception of the current scene but forgets the past; feeding history back into attention restores recall but dilutes the recent evidence that real-time perception depends on.
StreamTTT keeps the two demands in separate places.
Architecture
Every decoder layer of a pretrained Qwen3-VL-4B-Instruct carries two complementary memories:
- a sliding KV cache (4K tokens) inside the attention context, dedicated to recent evidence — this is the pretrained attention path, untouched;
- a parallel test-time-training (TTT) branch whose fast weights are updated online and store long-range history outside the attention context, so it never competes for context slots with recent tokens.
The TTT branch (FastWeightBlock) holds a small SwiGLU MLP (w0, w1, w2)
as its fast weights, updated chunkwise during the forward pass with
input-dependent, head-wise momentum and decay projected from the hidden states.
The two branch outputs are fused by a learnable channel-wise gate tanh(α),
initialized near zero so the model starts at the pretrained function and learns
to use long-term memory.
At inference, video is processed as ordered temporal windows: the KV cache is
pruned to its most recent L tokens after each window while the fixed-size TTT
state is carried forward without eviction, and M-RoPE positions stay globally
continuous across windows.
Key configuration (frozen into config.json):
| base model | Qwen/Qwen3-VL-4B-Instruct |
sliding_window |
4096 |
num_fw_heads / num_fw_kv_heads |
4 / 4 |
lact_chunk_size |
1024 |
lr_parameterization / ttt_base_lr |
ttt / 1e-4 |
ttt_momentum / ttt_weight_decay |
headwise / headwise |
use_residual / use_gate_for_memory |
True / True |
The fast-weight memory follows E2-TTT.
Results
Under each model's reported input protocol:
| Model | Input | StreamingBench RTVU | OVO-Bench RT | OVO-Bench BT | OVO-Bench Avg |
|---|---|---|---|---|---|
| HERMES-7B | 1 fps | 79.44 | 69.0 | 49.4 | 59.20 |
| SimpleStream-4B | 16 frames | – | 77.5 | 54.6 | 66.06 |
| SimpleStream-8B | 4 frames | 80.59 | 81.4 | 52.1 | 67.70 |
| StreamTTT-4B | 2 fps / 4K sliding KV cache | 81.32 | 78.1 | 59.9 | 69.00 |
We train StreamTTT jointly on offline long-video QA and a newly constructed real-time QA corpus. On OVO-Bench, under each model's reported input protocol, StreamTTT-4B outperforms the same-scale SimpleStream-4B by 0.6 points in real-time perception and 5.3 points in backward tracing. It also surpasses the larger SimpleStream-8B by 0.73 points on StreamingBench's Real-Time Visual Understanding (RTVU) subset.
Usage
from transformers import AutoModelForCausalLM, AutoProcessor
model = AutoModelForCausalLM.from_pretrained(
"zeyun-zhong/StreamTTT-4B", trust_remote_code=True, torch_dtype="bfloat16", device_map="cuda"
)
processor = AutoProcessor.from_pretrained("zeyun-zhong/StreamTTT-4B")
The processor ships with the checkpoint — do not substitute the base Qwen3-VL processor, whose video preprocessing config differs.
This gets you single-shot inference over a short clip. Streaming windowed
inference — carrying the TTT state across windows with globally continuous
M-RoPE — needs the helpers in streamttt/streaming/windowing.py; see the code
repository and docs/EVALUATION.md.
Intended use and limitations
Intended for research on streaming and long-form video understanding: real-time perception on live video, and backward tracing over history already seen.
The fixed-size TTT state is a lossy summary. It is not a substitute for full attention when the whole video fits in context. The paper (§4.3) quantifies this at a 4K budget against a 64K sliding-window reference. Use it where attention is genuinely bounded, not as a drop-in replacement.
No streaming free-form generation. The released evaluation path scores
multiple-choice answers by logits. OVO-Bench forward tasks (REC, SSR, CRR)
require free-form generation and raise NotImplementedError.
Domain. Training data is predominantly first- and third-person daily-activity video (egocentric recordings, YouTube activity clips, instructional video). Out-of-domain behaviour — surveillance, medical, sports broadcast, screen capture — is unverified.
Inherited limitations. Everything true of the Qwen3-VL-4B-Instruct base model (hallucination, OCR limits, language coverage) remains true here.
License
Model weights and code are Apache-2.0. The released annotations
(zeyun-zhong/RealTimeVideo-Instruct-112K)
are CC-BY-4.0, except adt_realtime_qa, which derives from Aria Digital Twin
ground truth and is CC-BY-NC-SA-4.0. No video is redistributed — each source
keeps its own terms, some requiring a signed agreement or barring commercial
use; see docs/DATA.md.
Citation
@article{chen2026streamttt,
title = {StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs},
author = {Chen, Joya and Zhong, Zeyun and Shou, Mike Zheng},
journal = {arXiv preprint arXiv:2608.13416},
year = {2026}
}
- Downloads last month
- 83
Model tree for zeyun-zhong/StreamTTT-4B
Base model
Qwen/Qwen3-VL-4B-Instruct