AI Tools.

Search

sentence similarity by cl-nagoya

ruri-v3-310m

Ruri v3 (310M) is Nagoya University's Japanese text embedding model built on the ModernBERT architecture, optimised for semantic similarity and retrieval in Japanese. It is part of the Ruri series, which targets Japanese-specific sentence embedding quality. The v3 310M variant balances embedding dimension, retrieval quality, and inference speed for production Japanese NLP pipelines.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository cl-nagoya/ruri-v3-310m at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
cl-nagoya
Pipeline tag
sentence-similarity
Weight formats
safetensors
License tag
apache-2.0 — read the license file in the repo before relying on it
Lineage
Language tags
Japanese (ja)
Papers cited
arXiv:2409.07737
Datasets declared
cl-nagoya/ruri-v3-dataset-ft
Downloads (HF counter at last fetch)
479,441
Likes (HF counter at last fetch)
82
Model card
https://huggingface.co/cl-nagoya/ruri-v3-310m

Use cases

  • Japanese semantic search and document retrieval
  • FAQ matching for Japanese customer service systems
  • Clustering Japanese text by topic or intent
  • Building Japanese RAG retrieval components
  • Evaluating ModernBERT effectiveness on Japanese language tasks

Pros

  • ModernBERT backbone improves on BERT for long-context Japanese text
  • Purpose-built for Japanese; outperforms multilingual embedding models on Japanese retrieval
  • 74 likes with Nagoya University academic backing
  • 310M provides more capacity than 100M-class Japanese embedding models

Cons

  • Japanese only; cannot be used for cross-lingual retrieval
  • No published JMTEB or JSQuAD benchmark comparison in the model card
  • ModernBERT architecture is relatively new; third-party tooling support may lag
  • v3 is a recent release; production stability should be verified with testing

Tags

safetensorsmodernbertsentence-similarityfeature-extractionjadataset:cl-nagoya/ruri-v3-dataset-ftarxiv:2409.07737base_model:cl-nagoya/ruri-v3-pt-310mbase_model:finetune:cl-nagoya/ruri-v3-pt-310mlicense:apache-2.0region:us