tt-tnt

tt-tnt is a 22M-parameter Llama-3-style base completion model with a 2048-token context.

  • It was trained from random initialization on one Tenstorrent Blackhole chip with ttml (tt-train), converted to Hugging Face format, and numerically verified against an independent reimplementation.
  • It is served on one Blackhole chip (mesh (1, 1)) through the Tenstorrent vLLM plugin, and it also runs on CPU with plain transformers.
  • The model demonstrates a pipeline; it isn't a capable model. It writes simple stories. It is not a chat model. The chat endpoint only continues the text you send.
  • Status: Functional: served and measured on hardware at concurrency 1 and 8, with no Model CI v0 run.
  • It was formerly published as tt-nanollama3 (see Lineage).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 6, v6 thin).

At a glance

Architecture Llama-3 style: RoPE (θ=500000), RMSNorm, SwiGLU, grouped-query attention. 22,025,088 parameters, hidden 384, 6 layers, 6 heads / 3 KV heads
Hardware One Blackhole chip, mesh (1, 1) (MESH_DEVICE=P150). Measured on one chip of a p300c in a TT-QuietBox 2. It hasn't been run on a physical P150 card
Context 2048 tokens (max_model_len 2048). Longer prompts get HTTP 400. Ignore the model_max_length sentinel in tokenizer_config.json
Vocabulary 32,000 (byte-level BPE, trained for this model)
Weights dtype bfloat16
License Apache-2.0 (weights and code). The training corpus includes share-alike sources; see Licensing
Status Functional: served on hardware, measured at concurrency 1 and 8
Model CI v0 not yet run

Intended use

Direct use:

  • A demonstration that a model can be designed, trained, packaged, and served entirely on Tenstorrent tooling, TT-native from the first line of code.
  • Pipeline, packaging, and serving experiments.
  • An instrument for small training studies.

Give it the opening of a simple story.

Out-of-scope use:

  • Questions, instructions, or conversation. It has no instruction tuning, and its training corpus has no dialogue or question-answer data at all.
  • Anything factual.
  • More than one chip: 3 KV heads don't divide any multi-chip mesh width.

Quickstart

uv tool install tenstorrent   # once — the Tenstorrent CLI, `tt`
tt model pull episod/tt-tnt
tt serve episod/tt-tnt

tt model pull installs the bundle into its own venv: pinned Python 3.12, ttnn==0.77.0, the empty-target vLLM 0.25.1 build, and the TT vLLM plugin. It downloads the weights by default, since tt has no --with-weights flag.

The server listens on port 20000 by default and walks upward if that port is busy. The first serve converts the weights into a tensor cache under the install folder. On a TT-QuietBox 2, that first serve reached a ready endpoint in 46 s, and a restart with the cache in place is faster. The server is ready when the log prints Application startup complete.

Without tt-cli, tt-model alone does the whole job:

tt-model pull  episod/tt-tnt --with-weights
tt-model serve episod/tt-tnt

Stop it with tt-model stop episod/tt-tnt. It opens one chip:

  • Under a chip grant (TT_VISIBLE_DEVICES, for example from a lease manager), run.sh uses the first device of the grant.
  • Otherwise it uses device 0.
  • On a two-chip p300c board, the bundled mesh-1x1.textproto lets the model open a single chip.

Serve profiles

profile hardware mesh max_num_seqs block_size max_model_len
P150 (default) one Blackhole chip (1, 1) 32 64 2048

This is the only profile. (1, 1) is the only mesh shape this model can serve at: the attention code requires both the head count (6) and the KV-head count (3) to divide the mesh width, and no multi-chip preset has width 3.

Using it

It is a base completion model, so the completions endpoint is the natural one:

curl -s http://localhost:20000/v1/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "episod/tt-tnt",
       "prompt": "Once upon a time, there was a little",
       "max_tokens": 128, "temperature": 0.8, "top_p": 0.95}'

/v1/chat/completions works, but the model does not chat. Since 2026-09-27 the tokenizer ships a plain chat template so that chat clients get a completion instead of an HTTP 400. The template works like this:

  • It invents no roles and drops system messages.
  • It joins the last 5 messages into one running text, and the model continues it.
  • Each line loses trailing spaces and tabs, blank lines are dropped, and lines are joined with a single space. That is how the training pipeline flattened its documents.

A user message is therefore a story opening, and the reply is its continuation:

curl -s http://localhost:20000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model": "episod/tt-tnt",
       "messages": [{"role": "user", "content": "Once upon a time, a little girl named Lily"}],
       "max_tokens": 80, "temperature": 0}'

Served on 2026-09-27 (greedy, 80 tokens), that request returned: " was very hungry. She wanted to eat some food, but her mom said she had to eat her favorite food. Lily was very hungry and wanted to eat her favorite food. …" Every prompt token vLLM received matched a local render of the template, in 3 of 3 requests, one of them a three-message history.

Sampling at temperature 0.8 / top_p 0.95 is the representative way to read this model. Greedy decoding produces repetition loops (see Limitations).

On CPU, no Tenstorrent hardware is required:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

tok = AutoTokenizer.from_pretrained("episod/tt-tnt")
model = AutoModelForCausalLM.from_pretrained("episod/tt-tnt").eval()

ids = tok("Once upon a time, there was a little", return_tensors="pt").input_ids
with torch.no_grad():
    out = model.generate(ids, max_new_tokens=60, do_sample=True, temperature=0.8, top_p=0.95)
print(tok.decode(out[0], skip_special_tokens=True))

Sample output

Greedy decoding, 60 new tokens, from the frozen evaluation set (docs/measurements/samples-tt-tnt-v3.md):

Once upon a time, there was a little girl named Lily. She loved to play outside in the park. One day, she saw a big, shiny rock on the ground. She picked it up and showed it to her mom. "Look, Mommy! I found a shiny rock!" she said. Her mom smiled and said, "That

And one that stops on its own rather than being cut off at the limit:

The ants had learned that being eaten was a way of helping others. The moral of the story is that it's important to be kind to others and to help others.

Both come from near the top of the frozen set of 15 and aren't typical of it. The linked file shows the full range, including the TinyStories collapses to "a little girl named Lily" and several hard repetition loops. Sampled output at temperature 0.8 is in docs/measurements/samples-tt-tnt-v3-t0.8.md and gives the more representative read of the model's range.

Expected performance

Serving throughput and latency (on device, 1 chip, measured 2026-09-27):

sampling ISL / OSL concurrency N TTFT median (p99) TPOT median (p99) tok/s per user output tok/s
greedy 128 / 128 1 16 5.11 ms (5.64) 1.81 ms (1.97) 551 539
greedy 128 / 128 8 64 14.53 ms (19.19) 2.00 ms (2.08) 501 3,836
greedy 1900 / 128 1 8 13.04 ms (14.38) 2.95 ms (3.00) 338 330
default (temperature 1.0) 128 / 128 1 16 6.48 ms (6.99) 3.45 ms (3.81) 290 286
default (temperature 1.0) 128 / 128 8 64 20.49 ms (26.06) 6.42 ms (6.61) 156 1,227

Methodology.

  • Hardware: one Blackhole chip (0000:01:00.0) of a p300c in a TT-QuietBox 2, mesh (1, 1).
  • Bundle: the 2026-09-27 v6 thin build, identical to this one except for the weights-revision string. Stack: ttnn 0.77.0, vLLM 0.25.1.
  • Tool: vllm bench serve over streaming /v1/completions, with random-token prompts (seed 0, --ignore-eos) and 2 warmup requests.
  • Each row is the second of two identical passes, so compile time is excluded.
  • "Default" sends no temperature, so vLLM's default of 1.0 applies. Most chat clients do the same.

What the numbers mean.

  • Sampling costs a lot on this stack. At concurrency 8, TPOT goes from 2.00 ms greedy to 6.42 ms with default sampling, and aggregate throughput drops about 3×. Send temperature: 0 where greedy output is acceptable.
  • The chip matters. The same bundle measured 2.53 ms c1 TPOT on another chip of the same box (0000:03:00.0) during tuning, against 1.77–1.81 ms on 0000:01:00.0. That wasn't investigated.
  • In the tuning sweep with the same settings, concurrency 32 reached 11,926 output tok/s and concurrency 64 reached 11,942, because the batch is capped at 32 (docs/measurements/serving-tune-2026-09-27.md).
  • An earlier batch-1 figure of 618 tok/s (2026-09-04) used whole-request wall time on a different bundle build. It isn't comparable to the per-token rates above.

Accuracy (CPU, fp32, 2048-token window, on the same weights, sha256 97e19118…):

metric value reference / chance source
HF conversion vs independent pure-NumPy ttml forward max abs logit diff ~6e-6 (correlation ≈ 1 − 1e-13) identical function docs/model-development-troubleshooting.md, tests/test_to_hf.py
StoryCloze (xstory_cloze en, 1,511 items) 0.5705 chance 0.5281 (measured class balance); context-blind 0.5222 README.md StoryCloze section, docs/measurements/storycloze-tt-tnt.json
wikitext bits/byte 1.4867 GPT-2 small 0.9769 lm-eval 0.4.9, docs/measurements/external-tt-tnt-v3.md
lambada_openai last-word acc 0.0798 GPT-2 small 0.3256 same
piqa acc / arc_easy acc 0.5326 / 0.2963 chance 0.50 / 0.25 same
hellaswag acc / arc_challenge acc / mmlu acc 0.2584 / 0.1817 / 0.2292 chance 0.25. At or below chance, so these aren't measurements of the model same
Terminates on </s> (frozen 15-prompt set, 128 max tokens) 5/15 greedy, 11/30 sampled previous checkpoint 0/15, 0/30 docs/measurements/behaviour-tt-tnt-v1-vs-v3.md

No on-device accuracy-vs-reference measurement (PCC or token agreement against CPU) has been recorded for tt-tnt.

Limitations

Context is 2048 tokens, enforced.

  • max_model_len is 2048.
  • A prompt of 2049 tokens gets a clean HTTP 400 ("maximum context length is 2048 tokens"), and the server keeps serving.
  • A 2032-token prompt succeeds.

max_num_seqs is 32, and that is a ceiling of the stack, not a tuning choice.

  • tt_transformers 0.77 supports batch sizes 1, 2, 4, 8, 16, and 32 only. With 64 or 128 the server fails to start (ValueError: Batch size 64 not supported). KV-cache capacity wasn't the limit: 133,120 tokens are allocated.
  • The only way past 32 is vLLM data parallelism. Two 1-chip replicas (DP2) were measured and lost at every concurrency up to 32: c1 TPOT 2.08 ms (+16%), c32 7,528 tok/s against 11,926. It was ahead only at concurrency 64 (12,854 against 11,942), and it costs a second chip, so DP2 isn't shipped. Details are in docs/measurements/serving-tune-2026-09-27.md.

It can't serve on more than one chip. With 3 KV heads, only a (1, 1) mesh works, so it can't use a P300 board's 2-chip or 4-chip meshes. It has been measured on one chip of a p300c, not on a physical P150 card. There is no Model CI v0 run.

The tokenizer is pinned; the on-device weights load from main.

  • The manifest pins weights.revision to the commit that added the chat template, and vLLM loads the tokenizer and template from that commit.
  • The tt_transformers model code loads the model config and weights from episod/tt-tnt without a revision, so it reads main. The two are identical today.
  • The local tensor cache is keyed by the commit it converted from. A card-only commit therefore triggers one re-conversion on the next serve, and stale weights are never served.

The headline validation loss isn't comparable to the previous checkpoint's. 2.9937 against 4.2203 looks like a large gain, but most of it isn't one.

  • The previous checkpoint's validation split was the tail 10% of a token stream whose sources are concatenated in sorted-name order. It therefore landed entirely inside wikipedia_simple.
  • By per-source measurement, that is the most out-of-domain source in the blend and the second-hardest: held-out loss 4.28, against TinyStories' 1.83. See docs/measurements/per-source-loss-tt-tnt-v1.md. That number measured domain transfer, not learning.
  • This run's split is stratified by source: a proportional tail from each of the nine sources. That makes it a fair sample of the training mixture, and a much easier one, because 31% of the mixture is TinyStories.
  • A meaningful share of the drop from 4.22 to 2.99 comes from the yardstick changing, not the model improving. Don't subtract the two numbers.

It has seen its training corpus once and hasn't memorized it.

  • At batch 16 and sequence length 2048, 10,764 steps is one epoch over the blend's 352.7M-token training split.
  • The validation curve (artifacts/checkpoints-tt-tnt-v3/val_losses.jsonl) falls from 5.084 at step 500 to ~3.28 by step 4,000.
  • After that it keeps drifting down slowly and noisily. The last ~2,300 steps oscillate between 2.87 and 3.10, against ~3.12 around steps 7,000–8,000.
  • Training hadn't clearly stopped improving when it ended. Still, the per-step gain over the last third is small enough that more steps at these settings wouldn't transform it.

It can now stop, which no previous checkpoint could.

  • Every earlier checkpoint was trained on a corpus containing zero </s> tokens, while its config.json declared eos_token_id: 2. Those checkpoints never ended generation on their own.
  • This one does: 5/15 greedy completions and 11/30 sampled completions end on </s>.
  • It's a partial fix. Two thirds of completions still run to the limit.
  • The model stops far more readily on short-document material than on the book-length sources. Separator density in the blend ranges from one per ~210 tokens in the short sources to one per ~80,000 in the books.

TinyStories still dominates its voice. The frozen evaluation set (docs/measurements/samples-tt-tnt-v3.md, greedy, 15 prompts) shows the pattern:

  • The model sometimes engages with a prompt's own material: sticks, roses, a procession, bees.
  • Several prompts still degenerate into hard repetition loops ("the rose is a rose, and the rose is a rose, and…").
  • The oblique, observational voice this blend targets is not present in this checkpoint.
  • Output under sampling is markedly better behaved.

Whether it uses all 2048 tokens is a separate question. The position-wise loss probe (docs/measurements/context-use-tt-tnt-v3.md) shows loss falling well past where the previous checkpoint went flat:

positions loss
[0,32) 4.23
[32,64) 3.28
[64,128) 2.94
[256,512) 2.85

Past ~256 tokens, though, the improvement is about 0.02 nats per bucket against a standard error of ~0.08. That is directionally right and inside the noise.

Its RMSNorm layers did learn.

  • The very first tt-tnt checkpoint trained with stochastic_rounding disabled. That silently froze all 13 RMSNorm gammas at 1.0, bfloat16's rounding fixed point.
  • This run set stochastic_rounding: true.
  • Read directly from the published weights, all 13 gammas have moved off 1.0 and none is degenerate. Per-tensor means range from 0.874 to 1.720, and per-tensor standard deviations from 0.034 to 0.207.

Risks and safety considerations

  • Repetition loops under greedy decoding. Use sampling (temperature 0.8, top_p 0.95), accepting the throughput cost shown above.
  • Not factual, and not conversational. It is a 22M-parameter story model with no instruction tuning. The chat endpoint returns a story continuation, not an answer. Anything that looks like a fact is a coincidence.
  • Conversion correctness is verified on CPU only. The HF export matches an independent NumPy forward to ~6e-6. No on-device accuracy-vs-reference measurement exists. On-device greedy output is fluent TinyStories-style prose.

Licensing

The model weights and this project's code are Apache-2.0.

The training corpus isn't. This checkpoint was trained on a nine-source blend. Two of those sources are share-alike:

Full per-source licence, attribution, and the pinned dataset revisions are in docs/corpus_licensing.md.

  • That file is generated from the project's source registry (train/corpus.py), not written by hand, so it can't drift from the registry.
  • Its "unsettled Data Derivative" language for share-alike sources applies to this checkpoint exactly as written there.

The corpus itself isn't redistributed. Only the recipe to reconstruct it byte-identically is published, as episod/tt-tnt-corpus: the source registry, pinned revisions, and fetch/prepare/measure/blend scripts.

Credit.

  • Architectural credit goes to Mini-LLM by Ashx098 for the component choices (RoPE, RMSNorm, SwiGLU, GQA, subword BPE). The originating lesson arc credits it. That repository declares no license, so this is a credit, not a license inheritance.
  • This implementation derives from tt-train's nanollama3 config and the ttml library.

Lineage

This model was originally published as tt-nanollama3.

  • It started as a nanollama3-like model: a Llama-3 architecture trained from random initialization with tt-train's ttml trainer, on TinyStories.
  • The architecture and trainer haven't changed. The config, train/configs/model/tt-tnt-384.yaml, is a verbatim copy of tt-train's own nanollama3.yaml.
  • What changed is the corpus and tokenizer, which the project now owns: a nine-source, licence-audited blend and a BPE tokenizer trained on that blend. That earned the new name, tt-tnt.

Three checkpoints have been published under this repo id:

corpus context document separators
the original (TinyStories-only) TinyStories alone 256 none
the first blend-trained checkpoint nine-source blend 512 none
this one nine-source blend, separator-carrying revision 2048 yes

Training (this checkpoint)

Corpus Nine-source, licence-audited blend: TinyStories, Simple English Wikipedia, and seven curated Project Gutenberg slices (docs/corpus_blend.md). This is the first revision to carry </s> separators, averaging one per ~478 tokens
Tokens seen 352,714,752: the full training split, one epoch
Steps 10,764 at batch 16, sequence length 2048 (32,768 tokens/step)
Hardware One Blackhole chip of a p300c (mesh_shape [1, 1]) in a TT-QuietBox 2. The other three chips were idle
Wall clock ~91 minutes
Final train loss 3.25
Final validation loss 2.9937 (end-of-run evaluate()). The periodic curve's last entry (step 10,764) reads 2.939; the two figures sample different held-out windows
Optimizer AdamW, constant lr 3e-4, weight decay 0.01, stochastic_rounding: true

Origin

This model grew out of the "Build an LLM from Scratch" lesson arc in tt-vscode-toolkit. It takes that arc past where the lessons stop: real training, checkpointing, conversion, and numerical verification.

  • The full build, including every dead end, is documented at tsingletaryTT/tt-tnt.
  • An earlier conversion loaded cleanly and generated fluent prose while computing the wrong function: a RoPE row-layout mismatch worth 1.3 nats. Only numerical comparison caught it. That story is in docs/model-development-troubleshooting.md.

Changelog

date change
2026-08-14 First tt-kernel vLLM bundle pushed; current weights published (separator-carrying blend, 2048 context, HF commit e166888600)
2026-08-15 Bundle re-pushed; model-card frontmatter restored after a tagging bug
2026-09-01 Bundle manifest republished at schema v5
2026-09-05 First v6 thin package
2026-09-15 Briefly a v5 fat (self-contained) package, then repackaged v6 thin (a69c5d8ba2); stale v5-fat tree removed
2026-09-27 Serving update; weights unchanged. The tokenizer gains a plain chat template, so /v1/chat/completions returns a continuation instead of HTTP 400. max_model_len is set to 2048, so over-length prompts get HTTP 400. The bundle ships mesh-1x1.textproto and uses the first chip of a grant, so it starts as one chip on a p300c. Adapter 1.1.0. The manifest records tt_metal_version 0.77.0 and pins the weights revision. Measured at concurrency 1 and 8, greedy and default sampling

The two earlier checkpoints (TinyStories-only and first blend) and the tt-nanollama3 → tt-tnt rename predate this repo's recorded history, and their dates aren't recorded here.

Related packages

  • episod/tt-tnt-1024 is the 123M sibling in the same from-scratch family. It is a different model, not a hardware variant:

    • hidden 1024, 8 layers, 4 KV heads, 512 context;
    • it adds a dialogue slice and 2.53B tokens of FineWeb-Edu;
    • its chat template answers in the Question: … Answer: … format it was trained on;
    • it serves on 4 chips (P300x2).

    Neither package supersedes the other. Losses aren't comparable across the two (2048 vs 512 windows).

  • episod/tt-tnt-corpus is the corpus recipe.

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/episod/tt-tnt/discussions. That is the one channel that reaches the bundle's author.

A problem with the tt tooling itself: use tt report issue. It collects your environment and opens a prefilled issue against tenstorrent/tt-cli, so it doesn't reach this package's author.

Product feedback: send it to support@tenstorrent.com.

Provenance

component built from
tt-metal ttnn==0.77.0 (PyPI pin in requirements.txt; the manifest records tt_metal_version: 0.77.0)
tt-metal models code tt-tnt-models-closure==0.77.0, a vendored models/common + models/tt_transformers/tt (not upstream tt-metal-models)
vLLM 0.25.1, empty target (wheels/vllm-0.25.1+empty-…whl)
vllm-tt-plugin 0.1.0 wheel built 2026-09-15 (sha256 3990ec4d…). Its source commit isn't recorded in the wheel
weights model.safetensors sha256 97e191180d8e845c…, first published in HF commit e1668886007997ad0c2195b208e9c5da9c0a48ce, = artifacts/hf-tt-tnt-v3. The manifest pins weights.revision to the commit that added the chat template
adapter tt_tnt_adapter.py 1.1.0 = bundle/tt_tnt_adapter.py at tsingletaryTT/tt-tnt commit 2830e43
mesh mesh-1x1.textproto = train/configs/mesh/mesh-1x1.textproto (tt-metal's p150 descriptor)
Model CI v0 not yet run
build 2026-09-27 · tt-model-manager 0.1.0 (tt_kernel_version), integration build of the v6 thin fixes · schema 6
Downloads last month
2,652
Safetensors
Model size
22M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train episod/tt-tnt