--- title: ModelLens emoji: ๐Ÿ”ญ colorFrom: indigo colorTo: pink sdk: gradio sdk_version: 4.44.0 python_version: '3.11' app_file: app.py pinned: false license: mit short_description: Finding the Best Model for Your Task from Myriads of Models --- # ModelLens โ€” Finding the Best Model for Your Task from Myriads of Models ๐Ÿ“„ **Paper**: [ModelLens: Finding the Best Model for Your Task from Myriads of Models](https://huggingface.co/papers/2605.07075)  ยท  ๐Ÿค— **Collection**: [luisrui/modellens](https://huggingface.co/collections/luisrui/modellens)  ยท  ๐Ÿ’ป **Code**: [github.com/luisrui/ModelLens](https://github.com/luisrui/ModelLens) Describe your dataset โ†’ pick a task and metric โ†’ get a ranked list of HuggingFace models likely to perform well on it. Backed by the ModelLens checkpoint trained on the cleaned + expanded `unified_augmented_v2` corpus, with a candidate pool of ~47k HuggingFace models. The full model uses learned model-id / model-description / dataset-id embeddings on top of the dataset-description and task/metric signals. ## How it works 1. Your dataset description is embedded with OpenAI `text-embedding-3-small` (1536-dim, the same encoder used during training). 2. The MLPMetric scores every candidate model conditioned on the embedding + chosen task + chosen metric. 3. We return the top-k, optionally filtered by parameter count, "official pretrained only", or "HuggingFace-hosted only". ## Bring your own OpenAI key This Space does **not** ship with a baked-in OpenAI key. Paste your own `sk-...` key into the "OpenAI API key" field โ€” it is sent directly to OpenAI for that single request and is **not stored, logged, or reused** by this Space. A query costs roughly **$0.000001** on your account (about a millionth of a dollar). If you don't have a key yet: https://platform.openai.com/api-keys ## Files in this Space ``` app.py Gradio entry point recommend.py Recommender (loads checkpoint + model pool, embeds dataset desc) inference_lib.py Self-contained MLPMetric implementation (no module/ tree needed) build_model_pool.py Offline helper to (re)build assets/model_pool.npz requirements.txt Pinned deps assets/ model_pool.npz Pre-computed candidate pool (47k models, size+family ids, popularity, HF urls) checkpoint/ MLPMetricFull.pt ~709 MB trained weights (slim: parent-class dead weights + train-set dataset_desc_matrix stripped) args.json Training-time hyperparameters (model dims, num_*) data/ task2id.json Task vocab metric2id.json Metric vocab ``` The Space looks for the checkpoint at `checkpoint/MLPMetricFull.pt` (or the legacy `checkpoint/MLPMetric.pt`) and the data JSONs at `data/`. Override with env vars `MODEL_CKPT`, `MODEL_ARGS`, `DATA_DIR`, `POOL_PATH` if you lay things out differently. ## Downloading the trained weights directly The `MLPMetricFull.pt` (~709 MB slim checkpoint) shipped in this Space *is* the released model โ€” you can pull it with `huggingface_hub`: ```python from huggingface_hub import hf_hub_download ckpt = hf_hub_download("spaces/luisrui/ModelLens", "checkpoint/MLPMetricFull.pt") args = hf_hub_download("spaces/luisrui/ModelLens", "checkpoint/args.json") ``` To load it without launching the Gradio app, see `inference_lib.py` and `recommend.py` in this repo; the slim checkpoint expects `strict=False` loading (3 inference-unused parent-class buffers are intentionally dropped โ€” see `slim_full_checkpoint.py`). ## Training corpus This model was trained on: - [`luisrui/ModelLens-corpus-v2`](https://huggingface.co/datasets/luisrui/ModelLens-corpus-v2) โ€” 1.81M rows, 47k models, expanded with HELM / LiveBench / OpenCompass (recommended) - [`luisrui/ModelLens-corpus-v1`](https://huggingface.co/datasets/luisrui/ModelLens-corpus-v1) โ€” 1.54M rows, 47k models, R1โ€“R6 cleaning pipeline applied (cleaner but smaller) ## Running locally ```bash cd web pip install -r requirements.txt # either set OPENAI_API_KEY in env, or paste it into the UI at runtime python app.py # open http://localhost:7860 ``` ## Rebuilding the model pool When you bump the candidate set (e.g. add new HF models to `model2id.json` / `model_profile.json`): ```bash python web/build_model_pool.py \ --data-dir data/unified_augmented_v2 \ --profile-dir data/unified_augmented \ --args checkpoint/mlp/unified_augmented_v2/FinalModel_v2_full_data_deployment/args.json \ --out web/assets/model_pool.npz \ --min-popularity 0 ``` (`--profile-dir` falls back to v1's `model_profile.json` / `model_popularity.json` for the ~21k v2 model names that v2 doesn't yet ship a profile for.)