Instructions to use bioinfoihb/Caduceus-fish20-adapted with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bioinfoihb/Caduceus-fish20-adapted with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="bioinfoihb/Caduceus-fish20-adapted", trust_remote_code=True)# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("bioinfoihb/Caduceus-fish20-adapted", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from bioinfoihb/Caduceus-fish20-adapted: direct link, hf CLI and curl.
- Browser
- Download file 1.2 kB
-
https://huggingface.co/bioinfoihb/Caduceus-fish20-adapted/resolve/main/README.md
- Command line
-
hf download hf://bioinfoihb/Caduceus-fish20-adapted/README.md
-
curl -L -o README.md https://huggingface.co/bioinfoihb/Caduceus-fish20-adapted/resolve/main/README.md
1.2 kB
metadata
language:
- en
tags:
- genomics
- dna
- masked-language-modeling
- caduceus
library_name: transformers
pipeline_tag: fill-mask
Caduceus-PS fish20 adapted
This checkpoint is Caduceus-PS adapted to the FishNALM 20-species fish-genome corpus for masked language modeling.
Domain-adaptive pretraining
| Setting | Value |
|---|---|
| Base model | Caduceus-PS (~7.7M parameters) |
| Tokenization | Single nucleotide |
| Hardware | 4 × NVIDIA A100 80GB |
| Per-device batch size | 64 |
| Gradient accumulation steps | 1 |
| Effective global batch size | 256 |
| Peak learning rate | 1 × 10⁻⁴ |
| Corpus exposure | 1 epoch |
Usage
This repository contains custom model and tokenizer code. Load it with trust_remote_code=True.
from transformers import AutoModelForMaskedLM, AutoTokenizer
model_id = "<your-namespace>/Caduceus-PS-fish20-adapted"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(model_id, trust_remote_code=True)
The checkpoint is intended for DNA-sequence representation and masked-token prediction. Evaluate and adapt it for each downstream task before use.