Instructions to use bioinfoihb/Caduceus-fish20-adapted with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bioinfoihb/Caduceus-fish20-adapted with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="bioinfoihb/Caduceus-fish20-adapted", trust_remote_code=True)# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("bioinfoihb/Caduceus-fish20-adapted", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Caduceus-PS fish20 adapted
This checkpoint is Caduceus-PS adapted to the FishNALM 20-species fish-genome corpus for masked language modeling.
Domain-adaptive pretraining
| Setting | Value |
|---|---|
| Base model | Caduceus-PS (~7.7M parameters) |
| Tokenization | Single nucleotide |
| Hardware | 4 × NVIDIA A100 80GB |
| Per-device batch size | 64 |
| Gradient accumulation steps | 1 |
| Effective global batch size | 256 |
| Peak learning rate | 1 × 10⁻⁴ |
| Corpus exposure | 1 epoch |
Usage
This repository contains custom model and tokenizer code. Load it with trust_remote_code=True.
from transformers import AutoModelForMaskedLM, AutoTokenizer
model_id = "<your-namespace>/Caduceus-PS-fish20-adapted"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(model_id, trust_remote_code=True)
The checkpoint is intended for DNA-sequence representation and masked-token prediction. Evaluate and adapt it for each downstream task before use.
- Downloads last month
- 26