Instructions to use facebook/MobileLLM-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use facebook/MobileLLM-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="facebook/MobileLLM-1B", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("facebook/MobileLLM-1B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use facebook/MobileLLM-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "facebook/MobileLLM-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "facebook/MobileLLM-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/facebook/MobileLLM-1B
- SGLang
How to use facebook/MobileLLM-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "facebook/MobileLLM-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "facebook/MobileLLM-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "facebook/MobileLLM-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "facebook/MobileLLM-1B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use facebook/MobileLLM-1B with Docker Model Runner:
docker model run hf.co/facebook/MobileLLM-1B
AttributeError: 'bool' object has no attribute 'vocab_size'
#2
by djstrong - opened
I can't test in lm-eval:
Traceback (most recent call last):
File "venv/bin/lm_eval", line 8, in <module>
sys.exit(cli_evaluate())
File "lm-evaluation-harness/lm_eval/__main__.py", line 369, in cli_evaluate
results = evaluator.simple_evaluate(
File "lm-evaluation-harness/lm_eval/utils.py", line 346, in _wrapper
return fn(*args, **kwargs)
File "lm-evaluation-harness/lm_eval/evaluator.py", line 192, in simple_evaluate
lm = lm_eval.api.registry.get_model(model).create_from_arg_string(
File "lm-evaluation-harness/lm_eval/api/model.py", line 148, in create_from_arg_string
return cls(**args, **args2)
File "lm-evaluation-harness/lm_eval/models/huggingface.py", line 254, in __init__
self.vocab_size = self.tokenizer.vocab_size
AttributeError: 'bool' object has no attribute 'vocab_size'
There's a typo in model card. Please use the following command to load tokenizer:
tokenizer=AutoTokenizer.from_pretrained("facebook/MobileLLM-1B", use_fast=False)
Alternatively, you can use lm-eval cli directly (for example arc_easy task):
lm_eval --model hf --model_args pretrained=facebook/MobileLLM-1B,trust_remote_code=True,use_fast_tokenizer=False --tasks arc_easy
zechunliu changed discussion status to closed