audio8-asr-infinite-p150
Streaming English speech recognition (Edge0/Audio8-ASR-Infinite: a Voxtral-style audio encoder feeding a Qwen2.5-3B-class decoder) on a single Blackhole chip, served as an HTTP file-transcription endpoint and a WebSocket streaming endpoint.
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Architecture | Mel features and two small causal convs run on the host; a 32-layer causal audio encoder (sliding window 750 frames), the projector and a 36-layer Qwen2 decoder run on the chip. The decoder's per-layer delay conditioning is folded into the norm weights at first start, and the decoder keeps a 30-second rolling context (upstream's own default) so streams can run for any length. |
| Hardware | p150 |
| License | apache-2.0 |
| Status | Beta. Built and measured on one chip of a p300c board; not validated on other hosts. Built from a local, unpushed tt-metal branch (audio8-asr) with uncommitted changes, so the recorded commit alone does not reproduce this image. |
Intended use
Direct use: Transcribing English speech, from files or live audio, with first text about 1.3 seconds after speech starts.
Out-of-scope use: Languages other than English, more than one simultaneous stream, speaker identification, and end-of-turn detection (the model's semantic-VAD heads are not run). Not evaluated on noisy, accented or far-field speech.
Quickstart
uv tool install tenstorrent # once β the Tenstorrent CLI, `tt`
tt model pull episod/audio8-asr-infinite-p150
tt serve episod/audio8-asr-infinite-p150
tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Edge0/Audio8-ASR-Infinite weights at 7476824bc222e4ad509d286e8cae8b8d3f371129 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Without tt-cli β tt-model alone does the whole job:
tt-model pull episod/audio8-asr-infinite-p150 --with-weights
tt-model serve episod/audio8-asr-infinite-p150
Using it
This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API β its request and response shapes are the model's own. See the author's notes above for the payload it expects.
File: curl -F file=@sample.wav -F response_format=verbose_json http://localhost:<port>/v1/audio/transcriptions. Live: WebSocket /v1/audio/stream; send binary frames of 16 kHz mono int16 PCM, receive {"type":"delta","text":...}, send finish for {"type":"final",...}. One stream at a time: a second is refused and an idle one is closed after 60 seconds. /health reports readiness.
Expected performance
Word error rate against ground truth on 73 LibriSpeech validation utterances (8.0 minutes, clean read speech) with a simple normalizer (digits and spelling variants are not merged, so figures are pessimistic): 7.30% when each utterance is sent separately; one continuous 8.6-minute stream: 5.22%. On the first 39 utterances the CPU reference implementation and this port both score 5.47%. Real time factor 0.66 to 0.68 on long streams (load average 4 to 6 from unrelated processes), first text 1.3 s after the stream starts, final text 0.9 s after the last audio. Details: doc/RESULTS.md. Against GPUs (indicative only; not like-for-like). Edge0 publishes no GPU latency, memory or throughput figure, only accuracy tables. The hardware numbers found are third-party, from one author, on an RTX 5090 Laptop GPU in bf16 at the 80 ms clock and 480 ms delay: the upstream torch decoder takes 20.1 ms per 80 ms step, the upstream vLLM server 21.8 ms at p50 (p99 73.1 ms; 71 of 7,653 steps in a 10-minute stream exceeded the 80 ms budget) and 12.9 GB of GPU memory; one fresh install on an RTX 4090 gave 16.6 ms per step for that author's own bf16 engine. This port averages about 54 ms per 80 ms step on one Blackhole chip (0.67 x 80 ms, end to end, including host work and the rolling re-prefill): roughly 2.5 times slower per step than the 5090 Laptop figures, but inside the real-time budget on average. Per-step tail latency (p99, max) and total device memory were not measured here. Accuracy is not comparable either: Edge0 reports 3.04% and 6.81% WER on LibriSpeech test-clean and test-other, and the third party reproduced 2.97% and 6.96% on 300 utterances; our 5.2% to 7.3% is on a different 73-utterance validation subset with a simpler normalizer. The like-for-like evidence we have is parity with the upstream CPU implementation on the same utterances (identical transcripts in 35 of 39, equal WER).
Limitations
English only; greedy decoding; one stream at a time; the 80 ms / 480 ms-delay profile only. The decoder's rolling context is rebuilt by re-prefilling the window, which differs from upstream's cache trimming (results above show it works). The MLP is zero-padded and prefill is chunked in 128 rows because other prefill shapes overflow L1. The first start compiles kernels and converts weights (71 seconds to ready on the development box with an empty cache, longer on a slower disk).
Risks and safety considerations
Output differs from the CPU reference on a few words (near-tied tokens under bf16/bfp8 arithmetic; 3 of 151 words on the first demo clips). The server binds all interfaces inside the container; publish its port only where you intend to. Audio is decoded with soundfile and resampled with scipy, not upstream's librosa path.
Licensing
Model weights and code: Apache-2.0 (Edge0/Audio8-ASR-Infinite). This package contains no weights; they are downloaded from the Hub at first use.
Related packages
Source code, tests, evaluation scripts and the Gradio demo: https://github.com/tsingletaryTT/tt-audio8-asr. Upstream implementation this port was validated against: https://github.com/Edge0-AI/Audio8-ASR-Infinite. Accuracy tables (Edge0): https://huggingface.co/Edge0/Audio8-ASR-Infinite. Third-party GPU measurements quoted above: https://github.com/Edge0-AI/Audio8-ASR-Infinite/issues/1 and https://github.com/scrappylabsai/audio8-asr-lean.
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/episod/audio8-asr-infinite-p150/discussions β that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from β code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | a local checkout β commit not published (dirty tree β the image includes uncommitted changes) |
code/ digest |
8d3856fc7656031f (sha256, first 16 hex digits) |
| built | 2026-10-01T18:02:43+00:00 by tt-model 0.1.0 |
Model tree for episod/audio8-asr-infinite-p150
Base model
Edge0/Audio8-ASR-Infinite