hikka-forge2vec

hikka-forge2vec is a unified multimodal embedding model for anime and manga, created by Lorg0n. It maps titles, descriptions, genres, metadata, and an optional poster to one L2-normalized 256-dimensional vector.

This repository is also an installable Python package. The public API treats embeddings as vector objects, so similarity search and vector arithmetic read like ordinary Python.

Similarity search in practice

The illustration below shows a full-profile query for Frieren: Beyond Journey's End: the nearby results share its fantasy-journey semantics, while the deliberately shown #59 and #60 neighbours have visibly lower cosine similarity. The score is a raw cosine value between L2-normalized vectors.

Forge2Vec similarity search illustration

Model details

  • Architecture: unified modality-token attention with a bounded visual gate
  • Embedding size: 256
  • Content types: anime and manga
  • Languages used during training: English, Ukrainian, and Japanese titles
  • Poster input: optional PIL.Image, NumPy array, or PyTorch tensor
  • Missing modalities: supported, including profiles without a poster
  • Visual influence: learned per item and bounded to at most 10%

The release contains the complete inference weights. It does not require a separate adapter checkpoint.

Conceptually, Forge2Vec reads descriptions, titles, aliases, genres, and metadata as distinct but related signals. Attention lets the model combine those signals while learning which parts of a profile matter for a particular work. A poster is encoded into a compact visual representation and passed through a learned per-item gate: the model can use visual style when it helps, but it cannot let a cover override the semantic profile. Missing fields are represented explicitly, so the same model also works with incomplete records. The fused result is normalized into one shared space for search and vector arithmetic.

Installation

Install directly from Hugging Face:

pip install git+https://huggingface.co/Lorg0n/hikka-forge2vec

Quick start

from PIL import Image
from hikka_forge import Forge2Vec, ForgeItem

forge = Forge2Vec()

frieren = ForgeItem(
    title="Frieren: Beyond Journey's End",
    native_title="Sousou no Frieren",
    synonyms=["Фрірен, що проводжає в останню путь"],
    synopsis="The elf mage Frieren begins another journey after defeating the Demon King.",
    synopsis_ua="Ельфійка-чарівниця Фрірен вирушає в нову подорож.",
    genres=["Adventure", "Drama", "Fantasy"],
    content_type="anime",
    year=2023,
    score=8.9,
    poster=Image.open("frieren.jpg"),
)

vector = forge.vec(frieren)
print(vector.shape)  # (256,)

The poster is optional:

vector = forge.vec(
    title="Frieren: Beyond Journey's End",
    synopsis="An immortal elf learns to understand time, memory, and people.",
    genres=["Adventure", "Drama", "Fantasy"],
    content_type="anime",
)

PyTorch and NumPy image values are accepted in CHW or HWC layout. Images may use the [0, 1], [0, 255], or already normalized [-1, 1] range:

import torch

poster = torch.rand(3, 720, 480)
vector = forge.vec(title="Example", poster=poster)

Catalogue search

Create an in-memory catalogue from anime and manga profiles:

catalogue = forge.catalogue(items, batch_size=32)

frieren = catalogue.vec("Frieren: Beyond Journey's End")
results = frieren.find(limit=10)

for match in results:
    print(match.rank, match.item.title, match.similarity)

You can also search with a new profile that is not already in the catalogue:

query = forge.vec(
    title="A quiet fantasy journey",
    synopsis="An old mage travels through a changing world.",
    genres=["Adventure", "Fantasy"],
)

results = catalogue.find(query, limit=10)

Vector arithmetic

ForgeVector supports addition, subtraction, scalar multiplication, division, and normalization:

frieren = catalogue.vec("Frieren: Beyond Journey's End")
fullmetal = catalogue.vec("Fullmetal Alchemist: Brotherhood")
attack_on_titan = catalogue.vec("Attack on Titan")

query = frieren - attack_on_titan + fullmetal
results = query.find(limit=10)

weighted_query = 0.7 * frieren + 0.3 * fullmetal
weighted_results = weighted_query.find(limit=10)

Catalogue vectors remember their catalogue. Works used to construct a query are excluded from its results automatically.

Input schema

Field Type Required Description
title str No Primary title
native_title str No Native-language title
synonyms Sequence[str] No Alternative titles
synopsis str No English or primary description
synopsis_ua str No Ukrainian description
genres Sequence[str] No Genre and theme labels
content_type "anime" or "manga" No Defaults to anime
year int No Release year
score float No Score on a 0-10 scale
poster PIL.Image, np.ndarray, or torch.Tensor No Poster image

At least one meaningful content field should be supplied. Unknown genres are handled as unknown rather than causing an error.

Training

The model was trained from anime and manga profiles derived from the hikka.io catalogue, together with a currently private training dataset.[^dataset-release] A profile combines multilingual and native titles, aliases, synopses, genres, production metadata, and—when available—a poster. The supervision includes recommendation-style similarity pairs, multilingual title variants, relationships between works, and synthetic examples for controlled vector arithmetic.

The difficult part was not simply making related works appear near one another. A model can do that while still placing a large cloud of merely genre-adjacent works at almost the same similarity. Generic fantasy wording, shared tags, common title fragments, and visually striking but irrelevant covers are all tempting shortcuts. Conversely, treating every non-positive pair as a hard negative can push out legitimate sequels, spin-offs, and works that share meaningful themes.

The final training procedure therefore used direct multi-task supervision and self-ranking: positive neighbours are kept ahead of selected alternatives, while distribution regularization discourages the embedding space from collapsing into a narrow high-cosine region. The negative curriculum combines mined hard cases with curated light negatives: candidates that look plausibly related from metadata or retrieval, but were judged unsuitable as recommendations. This gives the model a more useful boundary than random unrelated pairs alone, without teaching it that every imperfect match is wrong.

Poster style was first compressed through multi-level SigLIP 2 feature statistics. During unified training, a learned counterfactual gate estimates whether the poster improves the semantic representation for that individual work. The gate is bounded, so a visually memorable cover can refine a vector but cannot replace the work's text and metadata.

Base models and acknowledgements

Forge2Vec contains and builds upon two pretrained encoders. Both upstream models are distributed under the Apache License 2.0, which is also the license used for this release.

  • The text tower is based on Lorg0n/hikka-forge-paraphrase-multilingual-MiniLM-L12-v2, an anime-domain English/Ukrainian model fine-tuned by Lorg0n. That model is derived from sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2. Thanks to Nils Reimers, Iryna Gurevych, and the Sentence Transformers contributors for their work on Sentence-BERT and multilingual sentence embeddings.
  • The visual tower is based on google/siglip2-base-patch16-224. Thanks to Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner, Xiaohua Zhai, and the wider Google research team for SigLIP 2.
  • Thanks to the PyTorch, Hugging Face Transformers, Safetensors, and Hugging Face Hub contributors whose libraries make this release usable.

The complete Apache-2.0 terms are provided in LICENSE. Attribution for the bundled derivative weights is also recorded in NOTICE.

Evaluation

The current checkpoint was evaluated in the same mode in which the library is intended to be used: complete anime profiles with titles, descriptions, genres, metadata, and local posters. The catalogue contained 21,394 anime profiles. For each query, its own catalogue entry was excluded, all vectors were L2-normalized, and neighbours were ranked by raw cosine similarity.

Query Rank 1 Cosine Rank 2 Cosine
Frieren: Beyond Journey's End Frieren Season 2 0.9042 Kaina: Star Sage 0.8958
Fullmetal Alchemist: Brotherhood Sacred Star of Milos 0.8794 Fullmetal Alchemist 0.8605
Attack on Titan Attack on Titan Season 2 0.9353 Attack on Titan: Chronicle 0.9310
Overtake! Initial D Final Stage 0.7744 Geki Drive 0.7609
Berserk Golden Age Arc III 0.9462 Golden Age Arc Memorial Edition 0.9424

The same Frieren run also illustrates the long tail: rank 59 was Sunday Without God at 0.7284, and rank 60 was Chuumon no Ooi Ryouriten (1993) at 0.7273. These values are not probabilities or human relevance labels. They are distances in this model's learned space; the meaningful signal is the ordering and the separation between neighbours for the same query.

This is a qualitative retrieval check over a fixed catalogue, not an independent human-labelled leaderboard. It is intended to make the model's behaviour inspectable and reproducible, not to claim that every ranking is universally correct.

Limitations

  • Cosine similarity is a ranking score, not a calibrated probability.
  • Retrieval quality depends on the completeness and language of each profile.
  • The genre vocabulary is fixed; unknown labels map to an unknown category.
  • Posters have deliberately limited influence and may not change a result when the semantic profile is already confident.
  • Visual training targets broad style rather than literal pose or composition, but some visual shortcuts may remain.
  • The model reflects biases in the hikka.io catalogue and synthetic training supervision.
  • Vector arithmetic is an auxiliary capability and is less reliable than nearest-neighbour retrieval.

Citation

Forge2Vec

@misc{lorg0n2026forge2vec,
  author       = {Lorg0n},
  title        = {{hikka-forge2vec: Unified Multimodal Embeddings for Anime and Manga}},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Lorg0n/hikka-forge2vec}},
  note         = {Hugging Face model repository}
}

Sentence-BERT

@inproceedings{reimers2019sentencebert,
  title     = {Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
  author    = {Reimers, Nils and Gurevych, Iryna},
  booktitle = {Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing},
  year      = {2019},
  publisher = {Association for Computational Linguistics},
  url       = {https://arxiv.org/abs/1908.10084}
}

SigLIP 2

@misc{tschannen2025siglip2,
  title         = {SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
  author        = {Michael Tschannen and Alexey Gritsenko and Xiao Wang and Muhammad Ferjad Naeem and Ibrahim Alabdulmohsin and Nikhil Parthasarathy and Talfan Evans and Lucas Beyer and Ye Xia and Basil Mustafa and Olivier Hénaff and Jeremiah Harmsen and Andreas Steiner and Xiaohua Zhai},
  year          = {2025},
  eprint        = {2502.14786},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2502.14786}
}

[^dataset-release]: The training dataset is private for now. I hope this is temporary, and I plan to publish it in the future after it has been reviewed and prepared for a responsible release.

Downloads last month
13
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for Lorg0n/hikka-forge2vec