ANLP Assignment 1: Transformers from Scratch

Student: Yashav Bhatnagar
Roll number: 2024101030

This repository contains the trained checkpoints and preprocessing states for the five configurations used in ANLP Assignment 1.

Config Main setting
C1 Sinusoidal positions, MHA, LayerNorm, learned BPE
C2 RoPE, MHA, LayerNorm, learned BPE
C3 Sinusoidal positions, GQA, LayerNorm, learned BPE
C4 Sinusoidal positions, MHA, RMSNorm, learned BPE
C5 Sinusoidal positions, MHA, LayerNorm, entropy-based byte patches

Repository contents

  • C1/model.pt through C4/model.pt: best tokenized-model checkpoints
  • C1/src_tokenizer.pkl through C4/src_tokenizer.pkl: source BPE states
  • C1/tgt_tokenizer.pkl through C4/tgt_tokenizer.pkl: target BPE states
  • C5/model.pt: best C5 checkpoint
  • C5/patcher_state.json: fitted source and target entropy-patcher state

The models and BPE tokenizers were implemented from scratch in PyTorch for the course assignment. C5 groups every eight cipher bits into one byte before applying entropy-based dynamic patching.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support