A Controlled Study of Attention-Only Transformers
Paper • 2607.18363 • Published • 5
Simple Attention Network — a ~31M-parameter custom LLM, a faithful PyTorch port of the
needle architecture (arXiv:2607.18363). Trained locally on an RTX 4070 Ti 12 GB.
This repo contains the released weights for the 12-layer configuration.
| file | what |
|---|---|
pytorch_model.bin |
model state_dict (load via san_model.SimpleAttentionNetwork.load_state_dict) |
san_latest.pt |
raw training checkpoint (step 137,260): {step, loss, model_state_dict, optimizer_state_dict, config} |
config.json |
architecture hyperparameters |
tokenizer/ |
SmolLM2-135M tokenizer (vocab 49152) |
san_model.py, san_triton.py |
model definition (needed to load the weights) |
load_example.py |
minimal load + forward example |
from san_model import SimpleAttentionNetwork, SANConfig
import torch
cfg = SANConfig(num_layers=12)
model = SimpleAttentionNetwork(cfg).eval()
sd = torch.load("pytorch_model.bin", map_location="cpu")
model.load_state_dict(sd)
See load_example.py. Full training/eval code: GitHub
kenpeter/x-small.
MIT (code). Weights released for research use.