ANLP Assignment 1: Transformers from Scratch
Student: Yashav Bhatnagar
Roll number: 2024101030
This repository contains the trained checkpoints and preprocessing states for the five configurations used in ANLP Assignment 1.
| Config | Main setting |
|---|---|
| C1 | Sinusoidal positions, MHA, LayerNorm, learned BPE |
| C2 | RoPE, MHA, LayerNorm, learned BPE |
| C3 | Sinusoidal positions, GQA, LayerNorm, learned BPE |
| C4 | Sinusoidal positions, MHA, RMSNorm, learned BPE |
| C5 | Sinusoidal positions, MHA, LayerNorm, entropy-based byte patches |
Repository contents
C1/model.ptthroughC4/model.pt: best tokenized-model checkpointsC1/src_tokenizer.pklthroughC4/src_tokenizer.pkl: source BPE statesC1/tgt_tokenizer.pklthroughC4/tgt_tokenizer.pkl: target BPE statesC5/model.pt: best C5 checkpointC5/patcher_state.json: fitted source and target entropy-patcher state
The models and BPE tokenizers were implemented from scratch in PyTorch for the course assignment. C5 groups every eight cipher bits into one byte before applying entropy-based dynamic patching.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support