Models and Datasets of paper: [Simplified Sparse Attention via Gist Tokens]
Yuzhen Mao
gist-sparse-attention
·
AI & ML interests
None yet
Recent Activity
commentedon a paper 21 days ago
Simplified Sparse Attention via Gist Tokens commentedon a paper 21 days ago
Simplified Sparse Attention via Gist Tokens upvoted a paper 3 months ago
Simplified Sparse Attention via Gist TokensOrganizations
models 19
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk8
333k • Updated • 13
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk16
333k • Updated • 14
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk32
333k • Updated • 12
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4
333k • Updated • 69
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4
333k • Updated • 60
gist-sparse-attention/GSA-FT-Llama-3.2-1B-chunk16
1B • Updated • 15
gist-sparse-attention/GSA-FT-Llama-3.2-1B-chunk4-chunk4
1B • Updated • 14
gist-sparse-attention/GSA-link-FT-Llama-3.2-1B-chunk8
1B • Updated • 7
gist-sparse-attention/GSA-link-FT-Llama-3.2-1B-chunk16
1B • Updated • 22
gist-sparse-attention/GSA-link-FT-Llama-3.2-1B-chunk4-chunk4
1B • Updated • 7
datasets 14
gist-sparse-attention/GSA-FT-Llama-3.2-1B-data
Preview • Updated • 37
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data
Preview • Updated • 285
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data
Preview • Updated • 75
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk32-data
Viewer • Updated • 25.9k • 164
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk16-data
Preview • Updated • 107
gist-sparse-attention/GSA-FT-Qwen2-7B-Instruct-chunk8-data
Preview • Updated • 21
gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
Viewer • Updated • 88.6k • 136
gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
Viewer • Updated • 88.5k • 96
gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk32-data
Preview • Updated • 237
gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk16-data
Viewer • Updated • 88.8k • 236