Instructions to use arvindcodex11/MYRA-oss-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use arvindcodex11/MYRA-oss-1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="arvindcodex11/MYRA-oss-1", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("arvindcodex11/MYRA-oss-1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download benchmarks/evaluation.md from arvindcodex11/MYRA-oss-1: direct link, hf CLI and curl.
- Browser
- Download file 2.9 kB
-
https://huggingface.co/arvindcodex11/MYRA-oss-1/resolve/main/benchmarks/evaluation.md
- Command line
-
hf download hf://arvindcodex11/MYRA-oss-1/benchmarks/evaluation.md
-
curl -L -o evaluation.md https://huggingface.co/arvindcodex11/MYRA-oss-1/resolve/main/benchmarks/evaluation.md
evaluation protocol
indicemo
100 prompts span five equally represented delivery categories: happy, sad, angry, excited, and professional. each prompt combines two to five languages from english, hindi, telugu, tamil, kannada, bengali, and punjabi. five systems produce 500 recordings.
gemini 3.1 pro preview, qwen3.5-omni-plus, and gpt-realtime-2.1 score anonymized recordings on a 1-5 rubric covering tone, dynamics, phrasing, and sustained expression. aggregation uses the 98-prompt intersection with valid ratings from all judges: median per recording, mean within category, then an unweighted mean across categories. the emotions-only aggregate excludes professional.
rumik-oss-1 uses ira throughout; competing systems use language-conditioned voices. statistical significance of ranking differences has not been evaluated.
nova
seven systems are evaluated on 149 single-tag english prompts from nvv-superbench, with three synthesis runs per system. prompts use provider-specific tags. rumik-oss-1 uses ira, with happy descriptions for laugh/chuckle and sad descriptions for sigh.
silk-asr, scribe v2, and the nvv-superbench gemini verifier measure exact-position rendering. laugh and chuckle are pooled into a laughter category. the aggregate is the mean of laughter and sigh rendering rates, averaged across detectors. unsupported categories contribute zero, including cartesia's absent sigh tag. detector-level results report means and standard deviations across runs.
unrequested vocalization rates for rumik-oss-1 are 2.9% (silk-asr), 5.4% (gemini), and 4.5% (scribe). gemini is conditioned on the requested event; the transcribers receive audio only. silk-asr is developed by rumik; scribe shares a provider with an evaluated system.
wer and cer
the full comparison covers 15 languages. seven-language subsets are selected post hoc by top-two numerical rank among displayed systems. the wer subset includes malayalam in place of telugu. provider coverage varies by language; grok results are available only for hindi.
the wer protocol specifies 100 shared prompts per model and language, scored with indicconformer rnnt using normalized text. the cer evaluation reports 100 rumik recordings per language, except telugu with 99.
competitor sample counts, matched utterance coverage, normalization, and recognizer versions have not been independently verified. confidence intervals and significance tests are unavailable. wer and cer originate from separate result snapshots; matched recordings between them and checkpoint equivalence with the emotion/vocalization evaluations remain unverified.