StreamTTT-4B

Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs

📄 Paper · 💻 Code · 📊 Data

Streaming video assistants are pulled in two directions. Keeping only a short window of recent frames gives sharp perception of the current scene but forgets the past; feeding history back into attention restores recall but dilutes the recent evidence that real-time perception depends on.

StreamTTT keeps the two demands in separate places.

Architecture

Every decoder layer of a pretrained Qwen3-VL-4B-Instruct carries two complementary memories:

  • a sliding KV cache (4K tokens) inside the attention context, dedicated to recent evidence — this is the pretrained attention path, untouched;
  • a parallel test-time-training (TTT) branch whose fast weights are updated online and store long-range history outside the attention context, so it never competes for context slots with recent tokens.

The TTT branch (FastWeightBlock) holds a small SwiGLU MLP (w0, w1, w2) as its fast weights, updated chunkwise during the forward pass with input-dependent, head-wise momentum and decay projected from the hidden states. The two branch outputs are fused by a learnable channel-wise gate tanh(α), initialized near zero so the model starts at the pretrained function and learns to use long-term memory.

At inference, video is processed as ordered temporal windows: the KV cache is pruned to its most recent L tokens after each window while the fixed-size TTT state is carried forward without eviction, and M-RoPE positions stay globally continuous across windows.

Key configuration (frozen into config.json):

base model Qwen/Qwen3-VL-4B-Instruct
sliding_window 4096
num_fw_heads / num_fw_kv_heads 4 / 4
lact_chunk_size 1024
lr_parameterization / ttt_base_lr ttt / 1e-4
ttt_momentum / ttt_weight_decay headwise / headwise
use_residual / use_gate_for_memory True / True

The fast-weight memory follows E2-TTT.

Results

Under each model's reported input protocol:

Model Input StreamingBench RTVU OVO-Bench RT OVO-Bench BT OVO-Bench Avg
HERMES-7B 1 fps 79.44 69.0 49.4 59.20
SimpleStream-4B 16 frames – 77.5 54.6 66.06
SimpleStream-8B 4 frames 80.59 81.4 52.1 67.70
StreamTTT-4B 2 fps / 4K sliding KV cache 81.32 78.1 59.9 69.00

We train StreamTTT jointly on offline long-video QA and a newly constructed real-time QA corpus. On OVO-Bench, under each model's reported input protocol, StreamTTT-4B outperforms the same-scale SimpleStream-4B by 0.6 points in real-time perception and 5.3 points in backward tracing. It also surpasses the larger SimpleStream-8B by 0.73 points on StreamingBench's Real-Time Visual Understanding (RTVU) subset.

Usage

from transformers import AutoModelForCausalLM, AutoProcessor

model = AutoModelForCausalLM.from_pretrained(
    "zeyun-zhong/StreamTTT-4B", trust_remote_code=True, torch_dtype="bfloat16", device_map="cuda"
)
processor = AutoProcessor.from_pretrained("zeyun-zhong/StreamTTT-4B")

The processor ships with the checkpoint — do not substitute the base Qwen3-VL processor, whose video preprocessing config differs.

This gets you single-shot inference over a short clip. Streaming windowed inference — carrying the TTT state across windows with globally continuous M-RoPE — needs the helpers in streamttt/streaming/windowing.py; see the code repository and docs/EVALUATION.md.

Intended use and limitations

Intended for research on streaming and long-form video understanding: real-time perception on live video, and backward tracing over history already seen.

The fixed-size TTT state is a lossy summary. It is not a substitute for full attention when the whole video fits in context. The paper (§4.3) quantifies this at a 4K budget against a 64K sliding-window reference. Use it where attention is genuinely bounded, not as a drop-in replacement.

No streaming free-form generation. The released evaluation path scores multiple-choice answers by logits. OVO-Bench forward tasks (REC, SSR, CRR) require free-form generation and raise NotImplementedError.

Domain. Training data is predominantly first- and third-person daily-activity video (egocentric recordings, YouTube activity clips, instructional video). Out-of-domain behaviour — surveillance, medical, sports broadcast, screen capture — is unverified.

Inherited limitations. Everything true of the Qwen3-VL-4B-Instruct base model (hallucination, OCR limits, language coverage) remains true here.

License

Model weights and code are Apache-2.0. The released annotations (zeyun-zhong/RealTimeVideo-Instruct-112K) are CC-BY-4.0, except adt_realtime_qa, which derives from Aria Digital Twin ground truth and is CC-BY-NC-SA-4.0. No video is redistributed — each source keeps its own terms, some requiring a signed agreement or barring commercial use; see docs/DATA.md.

Citation

@article{chen2026streamttt,
  title   = {StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs},
  author  = {Chen, Joya and Zhong, Zeyun and Shou, Mike Zheng},
  journal = {arXiv preprint arXiv:2608.13416},
  year    = {2026}
}
Downloads last month
83
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for zeyun-zhong/StreamTTT-4B

Finetuned
(456)
this model

Papers for zeyun-zhong/StreamTTT-4B