YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Recurrent depth × context: causal-v4 exploratory experiment

Date: 2026-10-09 UTC

Question: Does recurrent depth amplify the benefit of longer context at tiny scale?

Reproducibility

Source: recurrent_ctx_interaction_v4_causal.py Source SHA-256: d5ef19a4cc0a6a26137865553e46a707882dae111ddd52832bd826f052dc0d09 Results SHA-256: 89acd1b3f005cdcf35f33a82bd1f194d9c3046a79c3531a7fdabd5b244b8ad75 Correctness preflight: passed (target alignment, causal attention, gradients, memory, checkpoint reload). Tracked execution: exit code 0; duration 61.64 seconds; no recorded stderr.

Byte-level TinyStories next-token prediction. A 95/5 contiguous raw-data split separates training from validation, with source-code caps of 50 million training bytes and 2 million validation bytes. Four block applications; d_model 256; four attention heads; vocabulary 256; tied token embeddings; AdamW; 2,000 updates; batch size 32; learning rate 0.0003; seed 42. Recurrent feed-forward width 2048; stacked width 128. Training uses random sampled sequences. Validation is performed in eval mode without gradients.

Recorded metrics

Configuration Parameters Context Training-token positions Final validation loss Final validation perplexity
recurrent_ctx64 1,397,504 64 4,096,000 1.80549 6.08295
recurrent_ctx256 1,446,656 256 16,384,000 2.23559 9.35204
stacked_ctx64 1,402,880 64 4,096,000 1.92167 6.83238
stacked_ctx256 1,452,032 256 16,384,000 2.27253 9.70391

Interpretation and limitations

Recurrent configurations achieved lower validation perplexity at both context lengths in this particular run. The recurrent-minus-stacked advantage in perplexity was approximately 0.749 at context 64 and 0.352 at context 256. This is exploratory evidence, not proof of an architecture benefit or a recurrence × context interaction. Context conditions processed different training-token counts (4.096M vs 16.384M) and were evaluated at different context lengths. Recurrent and stacked variants use different feed-forward widths to approximate parameter matching, not identical architectures. The random seed was initialized once, not reset per configuration. Validation means are unweighted averages of batch losses; the last partial batch can have disproportionate influence. Only one seed was tested. No inference-speed or broad quality evaluation is implied. Previous noncausal v3 results are invalid because they permitted future-token leakage.

A stronger follow-up would match training-token exposure, evaluate all conditions on an identical held-out target set using comparable available histories, use multiple independent seeds, and report confidence intervals.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support