audio8-asr-infinite-p150

Streaming English speech recognition (Edge0/Audio8-ASR-Infinite: a Voxtral-style audio encoder feeding a Qwen2.5-3B-class decoder) on a single Blackhole chip, served as an HTTP file-transcription endpoint and a WebSocket streaming endpoint.

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

At a glance

Architecture Mel features and two small causal convs run on the host; a 32-layer causal audio encoder (sliding window 750 frames), the projector and a 36-layer Qwen2 decoder run on the chip. The decoder's per-layer delay conditioning is folded into the norm weights at first start, and the decoder keeps a 30-second rolling context (upstream's own default) so streams can run for any length.
Hardware p150
License apache-2.0
Status Beta. Built and measured on one chip of a p300c board; not validated on other hosts. Built from a local, unpushed tt-metal branch (audio8-asr) with uncommitted changes, so the recorded commit alone does not reproduce this image.

Intended use

Direct use: Transcribing English speech, from files or live audio, with first text about 1.3 seconds after speech starts.

Out-of-scope use: Languages other than English, more than one simultaneous stream, speaker identification, and end-of-turn detection (the model's semantic-VAD heads are not run). Not evaluated on noisy, accented or far-field speech.

Quickstart

uv tool install tenstorrent   # once β€” the Tenstorrent CLI, `tt`
tt model pull episod/audio8-asr-infinite-p150
tt serve episod/audio8-asr-infinite-p150

tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Edge0/Audio8-ASR-Infinite weights at 7476824bc222e4ad509d286e8cae8b8d3f371129 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Without tt-cli β€” tt-model alone does the whole job:

tt-model pull  episod/audio8-asr-infinite-p150 --with-weights
tt-model serve episod/audio8-asr-infinite-p150

Using it

This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API β€” its request and response shapes are the model's own. See the author's notes above for the payload it expects.

File: curl -F file=@sample.wav -F response_format=verbose_json http://localhost:<port>/v1/audio/transcriptions. Live: WebSocket /v1/audio/stream; send binary frames of 16 kHz mono int16 PCM, receive {"type":"delta","text":...}, send finish for {"type":"final",...}. One stream at a time: a second is refused and an idle one is closed after 60 seconds. /health reports readiness.

Expected performance

Word error rate against ground truth on 73 LibriSpeech validation utterances (8.0 minutes, clean read speech) with a simple normalizer (digits and spelling variants are not merged, so figures are pessimistic): 7.30% when each utterance is sent separately; one continuous 8.6-minute stream: 5.22%. On the first 39 utterances the CPU reference implementation and this port both score 5.47%. Real time factor 0.66 to 0.68 on long streams (load average 4 to 6 from unrelated processes), first text 1.3 s after the stream starts, final text 0.9 s after the last audio. Details: doc/RESULTS.md. Against GPUs (indicative only; not like-for-like). Edge0 publishes no GPU latency, memory or throughput figure, only accuracy tables. The hardware numbers found are third-party, from one author, on an RTX 5090 Laptop GPU in bf16 at the 80 ms clock and 480 ms delay: the upstream torch decoder takes 20.1 ms per 80 ms step, the upstream vLLM server 21.8 ms at p50 (p99 73.1 ms; 71 of 7,653 steps in a 10-minute stream exceeded the 80 ms budget) and 12.9 GB of GPU memory; one fresh install on an RTX 4090 gave 16.6 ms per step for that author's own bf16 engine. This port averages about 54 ms per 80 ms step on one Blackhole chip (0.67 x 80 ms, end to end, including host work and the rolling re-prefill): roughly 2.5 times slower per step than the 5090 Laptop figures, but inside the real-time budget on average. Per-step tail latency (p99, max) and total device memory were not measured here. Accuracy is not comparable either: Edge0 reports 3.04% and 6.81% WER on LibriSpeech test-clean and test-other, and the third party reproduced 2.97% and 6.96% on 300 utterances; our 5.2% to 7.3% is on a different 73-utterance validation subset with a simpler normalizer. The like-for-like evidence we have is parity with the upstream CPU implementation on the same utterances (identical transcripts in 35 of 39, equal WER).

Limitations

English only; greedy decoding; one stream at a time; the 80 ms / 480 ms-delay profile only. The decoder's rolling context is rebuilt by re-prefilling the window, which differs from upstream's cache trimming (results above show it works). The MLP is zero-padded and prefill is chunked in 128 rows because other prefill shapes overflow L1. The first start compiles kernels and converts weights (71 seconds to ready on the development box with an empty cache, longer on a slower disk).

Risks and safety considerations

Output differs from the CPU reference on a few words (near-tied tokens under bf16/bfp8 arithmetic; 3 of 151 words on the first demo clips). The server binds all interfaces inside the container; publish its port only where you intend to. Audio is decoded with soundfile and resampled with scipy, not upstream's librosa path.

Licensing

Model weights and code: Apache-2.0 (Edge0/Audio8-ASR-Infinite). This package contains no weights; they are downloaded from the Hub at first use.

Related packages

Source code, tests, evaluation scripts and the Gradio demo: https://github.com/tsingletaryTT/tt-audio8-asr. Upstream implementation this port was validated against: https://github.com/Edge0-AI/Audio8-ASR-Infinite. Accuracy tables (Edge0): https://huggingface.co/Edge0/Audio8-ASR-Infinite. Third-party GPU measurements quoted above: https://github.com/Edge0-AI/Audio8-ASR-Infinite/issues/1 and https://github.com/scrappylabsai/audio8-asr-lean.

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/episod/audio8-asr-infinite-p150/discussions β€” that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

Provenance

The exact sources the image was built from β€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout β€” commit not published (dirty tree β€” the image includes uncommitted changes)
code/ digest 8d3856fc7656031f (sha256, first 16 hex digits)
built 2026-10-01T18:02:43+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for episod/audio8-asr-infinite-p150

Finetuned
(1)
this model