locate-anything-3b-p150
NVIDIA LocateAnything-3B (Eagle-family visual grounding / open-vocabulary detection VLM: MoonViT-SO-400M vision tower + Qwen2.5-3B-Instruct with a detection vocabulary) running entirely on one Tenstorrent Blackhole p150a via tt-nn: image + free-text query in, labelled boxes out. Weights: nvidia/LocateAnything-3B · Paper: arXiv:2605.27365 · Upstream code: NVlabs/Eagle (Embodied) · Port: changh95/tt-locate-anything
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart
tt-model pull changh95/locate-anything-3b-p150 --with-weights
tt-model serve changh95/locate-anything-3b-p150
- Weights
nvidia/LocateAnything-3Batc32291ca5e99go to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says
Application startup complete.
Run with tt-cli
tt serve changh95/locate-anything-3b-p150
{ printf '{"query":"car","image":"'; base64 -w0 media/demo_input.png; printf '"}'; } > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/locate-anything-3b-p150
POST /predict:image(base64 PNG/JPEG, one image),query(free text, 1-1000 chars; join categories with</c>, e.g.person</c>car); optionalmax_new_tokens(128, cap 1024),return_overlay(false).GET /health,GET /info.
Response
{"query": "car", "width": 1920, "height": 1080, "canonical_size": [616, 336], "grid_hw": [24, 44],
"raw_text": "<ref>car</ref><box><282><414><606><794></box><|im_end|>",
"detections": [{"label": "car", "box": [541.44, 447.12, 1163.52, 857.52], "box_norm": [282, 414, 606, 794]}],
"points": [], "num_generated_tokens": 10, "stopped_on_eos": true, "decode_mode": "ar_greedy_trace_device_sampling",
"timing_ms": {"vision": 1.4, "prefill": 86.3, "decode": 202.1, "total": 303.9, "decode_tok_s": 44.53}}
boxis[x1, y1, x2, y2]in original image pixels;box_normis the model's own 0..1000 output over the squashedcanonical_sizeview.pointsholds 2-coordinate outputs the same way.return_overlay: trueaddsoverlay_png_b64, the boxes drawn on your image as a base64 PNG.
Demo
Accuracy and speed
Current build (2026-10-03 optimization, code/ in this repo). Measured with code/locate_anything/tests/bench_pipeline.py (the same pipeline.run() path /predict uses: host preprocessing → upload → vision → prefill → greedy decode → text) on a Blackhole chip with a 12×10 compute grid and dispatch on ETH cores, warm, batch 1, demo image + car, 10 tokens; independently re-measured by a second run (interleaved with the previous build on the same chip).
| Metric | Current build (2026-10-03) | Previous release (2026-09-13, p150a) |
|---|---|---|
End-to-end pipeline.run() |
~228 ms median (bimodal 215 / 241 on a shared host) · 214 ms min | ~305 ms (/predict server-side median) |
| Vision trace (MoonViT, 1056 patches) | 26.8 ms | 48 ms (TT_FUSED=0) / submitted async in the fused graph |
| Prefill trace (+ first-token tail) | 36.0 ms (+2.3 ms) | 86 ms incl. the vision wait |
| Decode | 14.0 ms/token (~71 tok/s) · 9 steps 131 ms | 202 ms (~44.5 tok/s) |
| Host preprocessing · pixel upload | 19 ms · 1.5 ms | — |
Accuracy gate (test_decode_accuracy.py: vit_proj PCC > 0.99, teacher-forced step-logit PCC > 0.97, box count) |
PASS — vit_proj PCC 0.997; car: mean/min step PCC 0.9973/0.9933, top-1 10/10, box <282><414><607><797> = HF reference (IoU 1.000); wheel: 0.9966/0.9911, top-1 15/16, box IoU 0.761 / 0.981 (previous build 0.884 / 1.000: one near-tie coordinate token flips, <571> vs <550>) |
full-pipeline logits PCC 0.9919; box <282><414><606><794> (IoU 0.989 vs HF reference) |
RTX 5090 (same reference measurements as before, 2026-09-14: port's torch reference, eager PyTorch 2.11, batch 1, H2D/D2H included, plus the same host pre/post-processing; full table in GPU_COMPARISON.md):
| RTX 5090 precision | GPU served-like | vs current build (228 median / 214 min) | GPU vision+prefill vs ours (65 ms) | GPU 9-step decode vs ours (131 ms) |
|---|---|---|---|---|
| fp32 strict | 229.0 ms | parity (1.00× / Blackhole 1.07× faster at min) | 92.0 ms — Blackhole 1.41× faster | 123.5 ms — GPU 1.06× |
| fp32 + TF32 | 198.9 ms | GPU 1.15× / 1.08× | 61.9 ms — GPU 1.05× | 123.5 ms — GPU 1.06× |
| bf16 native weights | 171.6 ms | GPU 1.33× / 1.25× (previous release: 1.78×) | 49.1 ms — GPU 1.33× | 108.7 ms — GPU 1.21× |
bf16 native + torch.compile |
149.7 ms | GPU 1.52× / 1.43× (previous release: 2.04×) | 39.9 ms — GPU 1.63× | 96.3 ms — GPU 1.36× |
Caveats
- One image + one query per request, batch 1, requests are serialized. Every image is squash-resized onto a fixed 24×44-patch grid (616×336, 16:9;
LA_IN_TOKEN_LIMIT=1024, upstream default 25600): non-16:9 images are distorted before the model sees them, boxes still map back to original pixels. - bf16 vision, BF16 attention + BFP8 MLP LLM; the package ships an
mlp.pyoverlay (code/models/tt_transformers/tt/mlp.py, decodew1/w3spilled to DRAM) so the LLM fits L1 on one p150a. Fused device paths (MoonViT as one metal trace with the patch merger on device, device vision→LLM merge, traced prefill, one-row first token, on-device greedy argmax) are on by default since 2026-09-13;TT_FUSED=0restores the 2026-09-12 graph, and the vision numerics are unchanged bit for bit either way. Only greedy AR decode is served; the experimental Parallel Box Decoding (MTP) is not. - Weights are public and ungated but under NVIDIA's own license (not OSI);
7.7 GB download, and the first boot extracts a 6.8 GB Qwen2.5-3B checkpoint, converts it to BFP8 and JIT-compiles (100 s on the very first boot, ~50 s with cached weights but an empty kernel cache, ~17 s warm; measured 2026-09-13). - Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - Current build: tt-metal
8b98410e730pluspatches/tt-metal-eth-dispatch.patch(lets single-chip Blackhole open with dispatch on ETH cores, 1 CQ, freeing the Tensix dispatch column for a 12×10 compute grid); open the device withttnn.DispatchCoreConfig(ttnn.DispatchCoreType.ETH)(seecode/locate_anything/tests/bench_pipeline.py --dispatch eth). Previous release: tt-metalv0.78.0-dev20260820(main8b98410e730), p150a. - GPU comparison (current build): Blackhole matches fp32-strict RTX 5090 end to end and is 1.41× faster on vision + prefill; bf16 GPU is 1.33× faster overall (1.78× against the previous release), mostly in decode (71 vs 83 tok/s). RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128, no TensorRT / vLLM; medians of 50 iterations after warm-up, H2D/D2H included. Power was not measured on the Blackhole side, so no efficiency comparison is made. Full table:
GPU_COMPARISON.md.
Licensing
- Weights: nvidia/LocateAnything-3B,
other(NVIDIA license, not OSI); fetched from upstream, not redistributed here. - Port and serving code (
code/locate_anything,code/scripts): Apache-2.0 SPDX headers, from changh95/tt-locate-anything, published under the same upstream terms; the vendored tt-metal overlaycode/models/tt_transformers/tt/mlp.pyis Apache-2.0.
Provenance
The exact sources the container image was built from. code/ has since been updated (2026-10-03 optimized build, see OPT_REPORT.md) and is newer than the image; tt-model serve still runs the image's code until the image is rebuilt:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
27dc2b154aa86da9 (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-13T15:59:45+00:00 by tt-model 0.1.0 |

