vggt-1b-p150

Meta VGGT-1B (feed-forward multi-view 3D reconstruction) port on one Tenstorrent Blackhole p150a. Weights: facebook/VGGT-1B · Paper: arXiv:2503.11651 · Upstream code: facebookresearch/vggt · Port: changh95/tt-vggt

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a Python environment with tt-metal's ttnn, built at tt-metal 8b98410e730. ttnn is not on PyPI. patches/tt-metal-eth-dispatch.patch moves dispatch to ETH cores. On a p150a, this gives the 12×10 compute grid. from_pretrained and the server open the device with ETH dispatch by default. See Caveats.

hf download changh95/vggt-1b-p150 --exclude "image/*" --local-dir vggt-1b-p150 && cd vggt-1b-p150
pip install -e code/                   # adds torch, numpy<2, pillow, safetensors, huggingface_hub, einops, pyspng
pip install -e "code/[server,test]"    # also the HTTP server (fastapi, uvicorn, pydantic, zlib-ng, pybase64) and the tests (pytest, httpx, pyyaml)
from tt_vggt import VGGT

with VGGT.from_pretrained(device_id=0) as model:                  # weights from the HF cache, opens the chip, traces S=1..4, warms up
    out = model(["media/source_1.png", "media/source_2.png"])     # 1..4 views of ONE scene; view 0 is the world frame

print(out.depth.shape, out.extrinsic.shape)                       # (2, 518, 518, 1) (2, 3, 4)
points, colors = out.point_cloud(conf_threshold=1.5)              # (N, 3) float32 + (N, 3) uint8, padding excluded
Name Description
Input images 1 to 4 views of one scene, in order. Each view is a file path, PNG / JPEG bytes, a PIL.Image, or a numpy array / torch tensor (HxWx3 or 3xHxW, uint8 or float in 0..1). A stacked (S, H, W, 3), (S, 3, H, W) or (1, S, 3, H, W) array also works. Any size.
Option outputs=None "all" (default), "depth", "points", "camera", or a list of keys.
Option dtype="float32" Or "float16" (the server default) for depth, depth_conf and world_points_conf. world_points is always float32.
Option conf_threshold=None Pixels with a lower confidence become NaN in depth and world_points.
Option return_type="numpy", batch_dim=False "torch" gives torch tensors. batch_dim=True keeps the upstream leading batch axis of 1.
Output pose_enc, extrinsic, intrinsic (S, 9) upstream pose encoding; (S, 3, 4) OpenCV camera-from-world; (S, 3, 3) pinhole K in pixels of the 518×518 padded view.
Output depth, depth_conf (S, 518, 518, 1) depth along the camera z axis; (S, 518, 518) confidence (1 or more).
Output world_points, world_points_conf (S, 518, 518, 3) 3-D point of each pixel in the view-0 camera frame; (S, 518, 518) confidence.
Output images_u8, preprocess, timing_ms The preprocessed views, the resize / pad of each view, and the call timing.
Methods out.crop(key, i), out.valid_region(i), out.point_cloud(...) Remove the padding of view i; get the valid region; get the coloured points of all views.
  • from_pretrained warms up the model before it returns. The default warm-up traces S = 1, 2, 3 and 4 and runs 2 real calls for each S with outputs all, depth and points. With a warm kernel cache, this takes about 31-36 s. The first load on an empty kernel cache also compiles the kernels (some minutes).
  • After the warm-up, the first call is as fast as later calls. For the same input, the difference is 2 % or less. The time of a call also changes with the input image size (decode and resize).
  • model.warmup(S=..., outputs=..., dtype=..., conf_threshold=...) adds more variants. A variant that is already warm costs nothing.
  • The first from_pretrained downloads facebook/VGGT-1B (model.safetensors, 5.0 GB) to your HF cache.
  • Use with ... as model: or call model.close(). This releases the traces and closes the chip.
  • The paths media/... are relative. Run the snippet from the root of this repo, or give absolute paths.
  • Units: VGGT normalizes the scale of each scene. Depth, points and camera translations are not in metres.
  • model.map(scenes) gives the same results as one call per scene. It decodes the next scene on a host thread while the chip runs the current scene.
  • model(...) gives the same output bits as the HTTP POST /predict, including the NaN positions.
  • Full reference (all options, input forms, warm-up, tests): PYTHON.md. Runnable example: examples/quickstart.py. It runs from any directory and writes depth images and a coloured points.ply.

Serving (HTTP)

tt-model pull  changh95/vggt-1b-p150 --with-weights
tt-model serve changh95/vggt-1b-p150      # or, with tt-cli: tt serve changh95/vggt-1b-p150
printf '{"images":["%s"]}' "$(base64 -w0 media/input.png)" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/vggt-1b-p150
  • Weights facebook/VGGT-1B at 860abec7937d go to your HF cache; the image does not contain them.
  • Serves on port 20000 (or the next free port); ready when the log says Application startup complete.
  • POST /predict: images (1-4 base64 PNG/JPEG views of one scene, ordered; image is a single-view alias); optional output_format (npz | json), dtype (float16 | float32), outputs (subset of depth, depth_conf, world_points, world_points_conf), conf_threshold, json_stride (8). Also GET /health, GET /info.
  • The response has pose_enc, extrinsic, intrinsic, preprocess and timing_ms as JSON. dense.data is a base64 np.savez_compressed blob with depth (S,518,518), depth_conf, world_points (S,518,518,3) and world_points_conf. All pixel coordinates use the 518×518 padded view (x_518 = x_orig*scale_x + pad_left).
  • The container image runs the 2026-09-14 code (about 420 ms forward at S=1). The current code/ is faster. See Provenance.

Demo

Kitchen frames 00 / 03 (VGGT example scene) → the union of both views' world_points in the camera-1 frame, coloured by the source pixels, world_points_conf ≥ 1.5 (served npz output, code/make_demo.py).

Demo & Performances

Warm, batch 1, 518×518 views, all outputs, median, current code/. S is the number of views.

Metric Performance
Device trace (bench_vggt.py, S=1 / 2 / 3 / 4) 46.8 / 87.4 / 165.5 / 225.8 ms
Forward with upload and readback (bench_vggt.py served, S=1 / 2 / 3 / 4) 50.1 / 91.4 / 173.9 / 236.4 ms · S=1 min 49.2, S=2 min 90.4 · S=1 real photo 50.0
End-to-end /predict timing_ms.total (demo media, npz float16, S=1 / 2 / 3 / 4) 57.0 / 109.7 / 191.4 / 255.7 ms
Python model() call on the demo files (decode + forward + outputs, float32, S=1 / 2 / 3 / 4) 54.7 / 105.8 / 187.1 / 252.7 ms · forward part 50.7 / 92.8 / 174.8 / 239.4 ms
from_pretrained (warm kernel cache, default warm-up) 31-34 s · first call vs steady: +1.1 to +1.8 % for the worst of the 12 default variants

The measurement hardware is a Galaxy Blackhole chip in the p150a configuration: dispatch on ETH cores, 1 command queue and a 12×10 compute grid (the same grid as a p150a with ETH dispatch). The outputs are bit-identical to the earlier build that used worker dispatch with 2 command queues. The opt-in head variants (VGGT_FUSED_HEAD_VARIANTS=d,p, in Python head_variants=True) were not measured again in this configuration. See OPT_REPORT.md round 15. Accuracy against the fp32 torch reference: test_vggt.py S=1 min PCC 0.9971 (world_points_conf, gate 0.99), real-photo depth AbsRel mean 0.00480 (previous release 0.00479). Details: VERIFICATION_2026-10-04.md (with the 2026-10-05 ETH-dispatch check) and OPT_REPORT.md.

RTX 5090 reference measurements (2026-09-14) are unchanged. They use the port's torch reference in eager PyTorch 2.11, batch 1, with H2D/D2H included. The TT side is the bench_vggt.py served forward (50.1 / 91.4 ms, ETH dispatch, 1 command queue), which also includes upload and readback. Full table: GPU_COMPARISON.md.

RTX 5090 precision GPU S=1 vs current build, S=1 (50.1 ms) GPU S=2 vs current build, S=2 (91.4 ms)
fp32 strict 134.1 ms Blackhole 2.68× faster 245.0 ms Blackhole 2.68× faster
fp32 + TF32 93.0 ms Blackhole 1.86× faster 166.2 ms Blackhole 1.82× faster
bf16 autocast 65.6 ms Blackhole 1.31× faster 103.4 ms Blackhole 1.13× faster
fp16 autocast 63.3 ms Blackhole 1.26× faster 95.7 ms Blackhole 1.05× faster
bf16 weights (informational) 54.5 ms Blackhole 1.09× faster 84.6 ms GPU 1.08× faster
fp16 weights (informational) 52.3 ms Blackhole 1.04× faster 76.8 ms GPU 1.19× faster

The previous release was 6.4× (S=1) and 8.6× (S=2) slower than bf16 autocast GPU. Most of the gain comes from flash / custom SDPA attention, custom fused kernels for the DPT heads (upsample, conv and activation in one kernel) and the GELU epilogue, and one metal trace per S, with the readback of each output part between the trace parts.

Caveats

  • Does not scale to multiple p150a in a mesh configuration. The port is tuned for one chip with a 12×10 compute grid of Tensix cores. On a p150a, this grid needs the dispatch functions on ETH cores (patches/tt-metal-eth-dispatch.patch). Thus, this build assumes that you do not need chip-to-chip ethernet communication.
  • from_pretrained and the server open the chip with ETH dispatch, 1 command queue and a 12×10 grid (VGGT_DISPATCH=auto). If the chip does not open with ETH dispatch, they use stock dispatch with 1 command queue and show a RuntimeWarning. With stock dispatch, a p150a has an 11×10 grid. Start such a process with VGGT_GRID_CAP=0. The model then prints a warning, and the custom VSDPA attention kernel is replaced by stock SDPA. This fallback was tested only with a stub device. To use another device setup, open the device yourself and give it with device= (tt_vggt.device_open_kwargs() gives ETH dispatch, 1 command queue, the trace region and L1_SMALL).
  • VGGT_DISPATCH=worker (stock dispatch, with the option VGGT_NUM_CQ=2) is an opt-in for Galaxy Blackhole chips only. On a p150a, it gives an 11×10 grid.
  • Every view is resized and white-padded to 518×518. 1-4 ordered views of one scene per call, batch 1, calls are serialized on the one chip. Point tracking is not available (track head not ported).
  • bf16 on device with fp32 accumulation in the critical parts: outputs differ slightly from the fp32 reference. On the hard house_in_field photo, the single-image 0.99 PCC gate fails, as it also does on the previous release.
  • Weights are CC-BY-NC-4.0 (non-commercial only). The server and the API keep the torch model on the host (about 5 GB host RAM).
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150a power was not measured, so no efficiency comparison is made.

Licensing

  • Weights: facebook/VGGT-1B, CC-BY-NC-4.0 (non-commercial only; not redistributed here).
  • Port, Python API and serving code (code/models/, code/tt_vggt/, examples/): Apache-2.0 per changh95/tt-vggt, bound by the weights' non-commercial terms; vendored upstream code/vggt/ is Meta's under the VGGT License v1 (code/LICENSE.txt, with its Acceptable Use Policy).
  • patches/tt-metal-eth-dispatch.patch is a patch to tt-metal (Apache-2.0).

Provenance

These are the exact sources the container image was built from. code/ has since been updated (2026-10-04 optimized build, see OPT_REPORT.md; 2026-10-04 Python API, see PYTHON.md; 2026-10-05 ETH dispatch with 1 command queue as the default) and is newer than the image. tt-model serve runs the image's code until the image is rebuilt. tt-model.yaml and SERVING.md still describe the image:

component built from
tt-metal 8b98410e730bb504fea43a88609756e34821d91d
code/ digest (image) e847a1f38266f7c7 (sha256, first 16 hex digits; the current code/ differs)
built 2026-09-14T06:48:14+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/vggt-1b-p150

Base model

facebook/VGGT-1B
Finetuned
(11)
this model

Collection including changh95/vggt-1b-p150

Paper for changh95/vggt-1b-p150