vggt-1b-p150
Meta VGGT-1B (feed-forward multi-view 3D reconstruction) port on one Tenstorrent Blackhole p150a. Weights: facebook/VGGT-1B · Paper: arXiv:2503.11651 · Upstream code: facebookresearch/vggt · Port: changh95/tt-vggt
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a Python environment with tt-metal's ttnn, built at tt-metal 8b98410e730. ttnn is not on PyPI. patches/tt-metal-eth-dispatch.patch moves dispatch to ETH cores. On a p150a, this gives the 12×10 compute grid. from_pretrained and the server open the device with ETH dispatch by default. See Caveats.
hf download changh95/vggt-1b-p150 --exclude "image/*" --local-dir vggt-1b-p150 && cd vggt-1b-p150
pip install -e code/ # adds torch, numpy<2, pillow, safetensors, huggingface_hub, einops, pyspng
pip install -e "code/[server,test]" # also the HTTP server (fastapi, uvicorn, pydantic, zlib-ng, pybase64) and the tests (pytest, httpx, pyyaml)
from tt_vggt import VGGT
with VGGT.from_pretrained(device_id=0) as model: # weights from the HF cache, opens the chip, traces S=1..4, warms up
out = model(["media/source_1.png", "media/source_2.png"]) # 1..4 views of ONE scene; view 0 is the world frame
print(out.depth.shape, out.extrinsic.shape) # (2, 518, 518, 1) (2, 3, 4)
points, colors = out.point_cloud(conf_threshold=1.5) # (N, 3) float32 + (N, 3) uint8, padding excluded
| Name | Description | |
|---|---|---|
| Input | images |
1 to 4 views of one scene, in order. Each view is a file path, PNG / JPEG bytes, a PIL.Image, or a numpy array / torch tensor (HxWx3 or 3xHxW, uint8 or float in 0..1). A stacked (S, H, W, 3), (S, 3, H, W) or (1, S, 3, H, W) array also works. Any size. |
| Option | outputs=None |
"all" (default), "depth", "points", "camera", or a list of keys. |
| Option | dtype="float32" |
Or "float16" (the server default) for depth, depth_conf and world_points_conf. world_points is always float32. |
| Option | conf_threshold=None |
Pixels with a lower confidence become NaN in depth and world_points. |
| Option | return_type="numpy", batch_dim=False |
"torch" gives torch tensors. batch_dim=True keeps the upstream leading batch axis of 1. |
| Output | pose_enc, extrinsic, intrinsic |
(S, 9) upstream pose encoding; (S, 3, 4) OpenCV camera-from-world; (S, 3, 3) pinhole K in pixels of the 518×518 padded view. |
| Output | depth, depth_conf |
(S, 518, 518, 1) depth along the camera z axis; (S, 518, 518) confidence (1 or more). |
| Output | world_points, world_points_conf |
(S, 518, 518, 3) 3-D point of each pixel in the view-0 camera frame; (S, 518, 518) confidence. |
| Output | images_u8, preprocess, timing_ms |
The preprocessed views, the resize / pad of each view, and the call timing. |
| Methods | out.crop(key, i), out.valid_region(i), out.point_cloud(...) |
Remove the padding of view i; get the valid region; get the coloured points of all views. |
from_pretrainedwarms up the model before it returns. The default warm-up traces S = 1, 2, 3 and 4 and runs 2 real calls for each S with outputsall,depthandpoints. With a warm kernel cache, this takes about 31-36 s. The first load on an empty kernel cache also compiles the kernels (some minutes).- After the warm-up, the first call is as fast as later calls. For the same input, the difference is 2 % or less. The time of a call also changes with the input image size (decode and resize).
model.warmup(S=..., outputs=..., dtype=..., conf_threshold=...)adds more variants. A variant that is already warm costs nothing.- The first
from_pretraineddownloadsfacebook/VGGT-1B(model.safetensors, 5.0 GB) to your HF cache. - Use
with ... as model:or callmodel.close(). This releases the traces and closes the chip. - The paths
media/...are relative. Run the snippet from the root of this repo, or give absolute paths. - Units: VGGT normalizes the scale of each scene. Depth, points and camera translations are not in metres.
model.map(scenes)gives the same results as one call per scene. It decodes the next scene on a host thread while the chip runs the current scene.model(...)gives the same output bits as the HTTPPOST /predict, including the NaN positions.- Full reference (all options, input forms, warm-up, tests):
PYTHON.md. Runnable example:examples/quickstart.py. It runs from any directory and writes depth images and a colouredpoints.ply.
Serving (HTTP)
tt-model pull changh95/vggt-1b-p150 --with-weights
tt-model serve changh95/vggt-1b-p150 # or, with tt-cli: tt serve changh95/vggt-1b-p150
printf '{"images":["%s"]}' "$(base64 -w0 media/input.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/vggt-1b-p150
- Weights
facebook/VGGT-1Bat860abec7937dgo to your HF cache; the image does not contain them. - Serves on port 20000 (or the next free port); ready when the log says
Application startup complete. POST /predict:images(1-4 base64 PNG/JPEG views of one scene, ordered;imageis a single-view alias); optionaloutput_format(npz|json),dtype(float16|float32),outputs(subset ofdepth,depth_conf,world_points,world_points_conf),conf_threshold,json_stride(8). AlsoGET /health,GET /info.- The response has
pose_enc,extrinsic,intrinsic,preprocessandtiming_msas JSON.dense.datais a base64np.savez_compressedblob withdepth (S,518,518),depth_conf,world_points (S,518,518,3)andworld_points_conf. All pixel coordinates use the 518×518 padded view (x_518 = x_orig*scale_x + pad_left). - The container image runs the 2026-09-14 code (about 420 ms forward at S=1). The current
code/is faster. See Provenance.
Demo
Kitchen frames 00 / 03 (VGGT example scene) → the union of both views' world_points in the camera-1 frame, coloured by the source pixels, world_points_conf ≥ 1.5 (served npz output, code/make_demo.py).
Demo & Performances
Warm, batch 1, 518×518 views, all outputs, median, current code/. S is the number of views.
| Metric | Performance |
|---|---|
Device trace (bench_vggt.py, S=1 / 2 / 3 / 4) |
46.8 / 87.4 / 165.5 / 225.8 ms |
Forward with upload and readback (bench_vggt.py served, S=1 / 2 / 3 / 4) |
50.1 / 91.4 / 173.9 / 236.4 ms · S=1 min 49.2, S=2 min 90.4 · S=1 real photo 50.0 |
End-to-end /predict timing_ms.total (demo media, npz float16, S=1 / 2 / 3 / 4) |
57.0 / 109.7 / 191.4 / 255.7 ms |
Python model() call on the demo files (decode + forward + outputs, float32, S=1 / 2 / 3 / 4) |
54.7 / 105.8 / 187.1 / 252.7 ms · forward part 50.7 / 92.8 / 174.8 / 239.4 ms |
from_pretrained (warm kernel cache, default warm-up) |
31-34 s · first call vs steady: +1.1 to +1.8 % for the worst of the 12 default variants |
The measurement hardware is a Galaxy Blackhole chip in the p150a configuration: dispatch on ETH cores, 1 command queue and a 12×10 compute grid (the same grid as a p150a with ETH dispatch). The outputs are bit-identical to the earlier build that used worker dispatch with 2 command queues. The opt-in head variants (VGGT_FUSED_HEAD_VARIANTS=d,p, in Python head_variants=True) were not measured again in this configuration. See OPT_REPORT.md round 15. Accuracy against the fp32 torch reference: test_vggt.py S=1 min PCC 0.9971 (world_points_conf, gate 0.99), real-photo depth AbsRel mean 0.00480 (previous release 0.00479). Details: VERIFICATION_2026-10-04.md (with the 2026-10-05 ETH-dispatch check) and OPT_REPORT.md.
RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with H2D/D2H included. The TT side is the bench_vggt.py served forward (50.1 / 91.4 ms, ETH dispatch, 1 command queue), which also includes upload and readback. Full table: GPU_COMPARISON.md.
| RTX 5090 precision | GPU S=1 | vs current build, S=1 (50.1 ms) | GPU S=2 | vs current build, S=2 (91.4 ms) |
|---|---|---|---|---|
| fp32 strict | 134.1 ms | Blackhole 2.68× faster | 245.0 ms | Blackhole 2.68× faster |
| fp32 + TF32 | 93.0 ms | Blackhole 1.86× faster | 166.2 ms | Blackhole 1.82× faster |
| bf16 autocast | 65.6 ms | Blackhole 1.31× faster | 103.4 ms | Blackhole 1.13× faster |
| fp16 autocast | 63.3 ms | Blackhole 1.26× faster | 95.7 ms | Blackhole 1.05× faster |
| bf16 weights (informational) | 54.5 ms | Blackhole 1.09× faster | 84.6 ms | GPU 1.08× faster |
| fp16 weights (informational) | 52.3 ms | Blackhole 1.04× faster | 76.8 ms | GPU 1.19× faster |
The previous release was 6.4× (S=1) and 8.6× (S=2) slower than bf16 autocast GPU. Most of the gain comes from flash / custom SDPA attention, custom fused kernels for the DPT heads (upsample, conv and activation in one kernel) and the GELU epilogue, and one metal trace per S, with the readback of each output part between the trace parts.
Caveats
- Does not scale to multiple p150a in a mesh configuration. The port is tuned for one chip with a 12×10 compute grid of Tensix cores. On a p150a, this grid needs the dispatch functions on ETH cores (
patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication. from_pretrainedand the server open the chip with ETH dispatch, 1 command queue and a 12×10 grid (VGGT_DISPATCH=auto). If the chip does not open with ETH dispatch, they use stock dispatch with 1 command queue and show aRuntimeWarning. With stock dispatch, a p150a has an 11×10 grid. Start such a process withVGGT_GRID_CAP=0. The model then prints a warning, and the custom VSDPA attention kernel is replaced by stock SDPA. This fallback was tested only with a stub device. To use another device setup, open the device yourself and give it withdevice=(tt_vggt.device_open_kwargs()gives ETH dispatch, 1 command queue, the trace region and L1_SMALL).VGGT_DISPATCH=worker(stock dispatch, with the optionVGGT_NUM_CQ=2) is an opt-in for Galaxy Blackhole chips only. On a p150a, it gives an 11×10 grid.- Every view is resized and white-padded to 518×518. 1-4 ordered views of one scene per call, batch 1, calls are serialized on the one chip. Point tracking is not available (track head not ported).
- bf16 on device with fp32 accumulation in the critical parts: outputs differ slightly from the fp32 reference. On the hard
house_in_fieldphoto, the single-image 0.99 PCC gate fails, as it also does on the previous release. - Weights are CC-BY-NC-4.0 (non-commercial only). The server and the API keep the torch model on the host (about 5 GB host RAM).
- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150a power was not measured, so no efficiency comparison is made.
Licensing
- Weights: facebook/VGGT-1B, CC-BY-NC-4.0 (non-commercial only; not redistributed here).
- Port, Python API and serving code (
code/models/,code/tt_vggt/,examples/): Apache-2.0 per changh95/tt-vggt, bound by the weights' non-commercial terms; vendored upstreamcode/vggt/is Meta's under the VGGT License v1 (code/LICENSE.txt, with its Acceptable Use Policy). patches/tt-metal-eth-dispatch.patchis a patch to tt-metal (Apache-2.0).
Provenance
These are the exact sources the container image was built from. code/ has since been updated (2026-10-04 optimized build, see OPT_REPORT.md; 2026-10-04 Python API, see PYTHON.md; 2026-10-05 ETH dispatch with 1 command queue as the default) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:
| component | built from |
|---|---|
| tt-metal | 8b98410e730bb504fea43a88609756e34821d91d |
code/ digest (image) |
e847a1f38266f7c7 (sha256, first 16 hex digits; the current code/ differs) |
| built | 2026-09-14T06:48:14+00:00 by tt-model 0.1.0 |
Model tree for changh95/vggt-1b-p150
Base model
facebook/VGGT-1B