AI Tools.

Search

sentence similarity by LazarusNLP

all-indo-e5-small-v4

all-indo-e5-small is LazarusNLP's Indonesian fine-tune of a small e5 embedding model, designed to improve semantic search and sentence similarity quality on Bahasa Indonesia text. v4 reflects iterative improvements over previous Indonesian embedding baselines.

Summary text generated by an automated pipeline from the model card · Not individually reviewed or run by us · How this page is made

From the model card

Fields below are copied from the tags and counters on the HuggingFace repository LazarusNLP/all-indo-e5-small-v4 at our last fetch. They are set by the uploader, not verified by us; rows with no tag are omitted. How this page is made.

Publisher (HF namespace)
LazarusNLP
Pipeline tag
sentence-similarity
Library
Sentence Transformers, Transformers
Weight formats
ONNX, safetensors
Datasets declared
indonli, indolem/indo_story_cloze, unicamp-dl/mmarco, miracl/miracl, nthakur/swim-ir-monolingual, LazarusNLP/multilingual-NLI-26lang-2mil7-id, SEACrowd/wrete, SEACrowd/indolem_ntp and 5 more on the model card
Downloads (HF counter at last fetch)
357,255
Likes (HF counter at last fetch)
13
Model card
https://huggingface.co/LazarusNLP/all-indo-e5-small-v4

Use cases

  • Semantic search over Indonesian-language document collections
  • Indonesian-language FAQ retrieval for chatbot grounding
  • Clustering Indonesian news articles by topic
  • Embedding Indonesian social media text for similarity tasks

Pros

  • Specific fine-tuning on Indonesian text outperforms generic multilingual embedders on Bahasa
  • Small model size keeps inference cost low
  • Fills a real gap — Indonesian NLP resources are sparser than major languages
  • Iterative v4 release suggests ongoing quality improvements

Cons

  • Primarily Bahasa Indonesia standard — Javanese, Sundanese dialects not covered
  • No published MTEB Indonesian subset scores to compare against alternatives
  • Small model limits embedding quality on long passages
  • Community project without commercial support

Tags

sentence-transformersonnxsafetensorsbertfeature-extractionsentence-similaritytransformersdataset:indonlidataset:indolem/indo_story_clozedataset:unicamp-dl/mmarcodataset:miracl/miracldataset:nthakur/swim-ir-monolingualdataset:LazarusNLP/multilingual-NLI-26lang-2mil7-iddataset:SEACrowd/wretedataset:SEACrowd/indolem_ntpdataset:khalidalt/tydiqa-goldpdataset:SEACrowd/facqadataset:indonesian-nlp/lfqa_iddataset:jakartaresearch/indoqadataset:jakartaresearch/id-paraphrase-detection