Sentence Similarity
sentence-transformers
Safetensors
bert
feature-extraction
Generated from Trainer
dataset_size:115403
loss:MultipleNegativesSymmetricRankingLoss
loss:TripletLoss
loss:CosineSimilarityLoss
loss:CoSENTLoss
Eval Results (legacy)
text-embeddings-inference
Instructions to use jangedoo/all-MiniLM-L6-v3-nepali with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use jangedoo/all-MiniLM-L6-v3-nepali with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("jangedoo/all-MiniLM-L6-v3-nepali") sentences = [ "प्रहरीको पोसाक लगाएर एक समूहले गिरफ्तार गरिरहेका छन् ।", "एउटा मानिस अण्डाको बोट काट्दै छ।", "चार कुकुरहरू हिउँमा खेल्छन् र तिनीहरूको पछाडि शहरको स्काइलाइन।", "प्रहरी अधिकारीहरूको एक समूह सुरक्षा लगाएको छ।" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
SentenceTransformer based on jangedoo/all-MiniLM-L6-v2-nepali
This is a sentence-transformers model finetuned from jangedoo/all-MiniLM-L6-v2-nepali on the title_excerpt, ne_en, excerpt_paraphrase, nepali_triplets, stsb_en, stsb_ne, stsb_en_ne and stsb_ne_en datasets. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: jangedoo/all-MiniLM-L6-v2-nepali
- Maximum Sequence Length: 256 tokens
- Output Dimensionality: 384 dimensions
- Similarity Function: Cosine Similarity
- Training Datasets:
- title_excerpt
- ne_en
- excerpt_paraphrase
- nepali_triplets
- stsb_en
- stsb_ne
- stsb_en_ne
- stsb_ne_en
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'max_seq_length': 256, 'do_lower_case': False}) with Transformer model: BertModel
(1): Pooling({'word_embedding_dimension': 384, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
(2): Normalize()
)
Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("jangedoo/all-MiniLM-L6-v3-nepali")
# Run inference
sentences = [
'कालोपत्रे भएपछि यस्तो बन्यो थानकोटको फ्लाइओभर (तस्वीरहरू)',
'Thankot flyover looks like this after being tarred (photos)',
'58 hybrid vehicles entered in 10 months',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [3, 3]
Evaluation
Metrics
Information Retrieval
- Datasets:
multi_lang_ir,en_irandne_ir - Evaluated with
InformationRetrievalEvaluator
| Metric | multi_lang_ir | en_ir | ne_ir |
|---|---|---|---|
| cosine_accuracy@10 | 0.9174 | 0.9606 | 0.8874 |
| cosine_precision@10 | 0.0917 | 0.0961 | 0.0887 |
| cosine_precision@50 | 0.0193 | 0.0197 | 0.0191 |
| cosine_recall@10 | 0.9174 | 0.9606 | 0.8874 |
| cosine_recall@50 | 0.967 | 0.9871 | 0.9532 |
| cosine_ndcg@10 | 0.8412 | 0.9163 | 0.7886 |
| cosine_mrr@10 | 0.8165 | 0.9017 | 0.7568 |
| cosine_map@100 | 0.8192 | 0.9032 | 0.7604 |
Translation
- Dataset:
translation - Evaluated with
TranslationEvaluator
| Metric | Value |
|---|---|
| src2trg_accuracy | 0.5392 |
| trg2src_accuracy | 0.575 |
| mean_accuracy | 0.5571 |
Triplet
- Dataset:
nepali_triplets - Evaluated with
TripletEvaluatorwith these parameters:{ "margin": { "cosine": 0.1, "dot": 0.1, "manhattan": 0.1, "euclidean": 0.1 } }
| Metric | Value |
|---|---|
| cosine_accuracy | 0.54 |
Semantic Similarity
- Datasets:
stsb_enandstsb_ne - Evaluated with
EmbeddingSimilarityEvaluator
| Metric | stsb_en | stsb_ne |
|---|---|---|
| pearson_cosine | 0.8818 | 0.6782 |
| spearman_cosine | 0.8791 | 0.6788 |
Training Details
Training Datasets
title_excerpt
title_excerpt
- Dataset: title_excerpt
- Size: 86,649 training samples
- Columns:
titleandexcerpt - Approximate statistics based on the first 1000 samples:
title excerpt type string string details - min: 7 tokens
- mean: 28.1 tokens
- max: 92 tokens
- min: 24 tokens
- mean: 73.46 tokens
- max: 256 tokens
- Samples:
title excerpt Rare Cheer Pheasant population rises in Kaligandaki BasinThe population of the rare Cheer Pheasant has increased in the Kaligandaki River basin areas of Myagdi and Mustang, attributed to conservation efforts and anti-poaching measures.विदेश पठाउने भन्दै १४ करोड १३ लाख ठगी, एकजना पक्राउवैदेशिक रोजगारीमा पठाउने भन्दै १४ करोड १३ लाख रुपैयाँ ठगी गरेको आरोपमा मकवानपुरका ५२ वर्षीय शम्भु दयाल अग्रवाल पक्राउ परेका छन्।Longstanding road encroachment issue in Hetauda resolvedA major demolition drive in Hetauda cleared 530 illegal structures encroaching on key highways to reclaim road boundaries after years of legal disputes. - Loss:
MultipleNegativesSymmetricRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim" }
ne_en
ne_en
- Dataset: ne_en
- Size: 3,765 training samples
- Columns:
titleandtranslation - Approximate statistics based on the first 1000 samples:
title translation type string string details - min: 4 tokens
- mean: 30.39 tokens
- max: 84 tokens
- min: 3 tokens
- mean: 25.22 tokens
- max: 78 tokens
- Samples:
title translation FWLD pushes for inclusive, non-discriminatory citizenship lawFWLD समावेशी, गैर-भेदभावरहित नागरिकता कानूनको लागि जोड दिन्छ
Clouded leopard: A vanishing jewel of the forest
बादलयुक्त चितुवा: जंगलको हराउने रत्नसिटिजन्स सदाबहार इकाइमा आवेदन म्याद थपApplication deadline extended to Citizens Sadabahar Unit - Loss:
MultipleNegativesSymmetricRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim" }
excerpt_paraphrase
excerpt_paraphrase
- Dataset: excerpt_paraphrase
- Size: 549 training samples
- Columns:
sentence1andsentence2 - Approximate statistics based on the first 549 samples:
sentence1 sentence2 type string string details - min: 22 tokens
- mean: 77.0 tokens
- max: 180 tokens
- min: 26 tokens
- mean: 82.48 tokens
- max: 184 tokens
- Samples:
sentence1 sentence2 भरतपुर कारागारमा मौसम परिवर्तन र वर्षातका कारण २२ कैदीबन्दीलाई रुघा र ज्वरोको समस्या देखिएको छ र उनीहरूको छुट्टै उपचार भैरहेको छ।मौसममा आएको बदलाव र वर्षाका कारण भरतपुर जेलमा २२ जना कैदीबन्दीलाई रुघाखोकी र ज्वरोले सताएको छ, जसका लागि उनीहरूको विशेष उपचार भइरहेको छ।Heavy to very heavy rainfall is likely today in six provinces of Nepal due to the monsoon trough near its average position, with thunder and lightning expected in several regions.Intense precipitation is anticipated this day across six Nepalese provinces, owing to the monsoon trough maintaining its usual course, and electrical storms are predicted in various areas.भोजपुरको रामप्रसाद राई गाउँपालिका–६ बैकुण्ठेका पाँच सामुदायिक विद्यालयका ४३८ विद्यार्थीलाई पोसाक वितरण गरिएको छ। वडाले विद्यालयबीच एकरुपता ल्याउने लक्ष्यले पोसाक, जुत्ता र टिसर्ट वितरण गरेको हो।रामप्रसाद राई गाउँपालिका–६, भोजपुरको बैकुण्ठेस्थित पाँच सामुदायिक विद्यालयका ४३८ जना विद्यार्थीहरूलाई वडाले विद्यालयहरूमा एकरूपता कायम गर्नका लागि पोसाक, जुत्ता तथा टिसर्ट प्रदान गरेको छ। - Loss:
MultipleNegativesSymmetricRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim" }
nepali_triplets
nepali_triplets
- Dataset: nepali_triplets
- Size: 1,600 training samples
- Columns:
sentence,positive_sentence, andnegative_sentence - Approximate statistics based on the first 1000 samples:
sentence positive_sentence negative_sentence type string string string details - min: 25 tokens
- mean: 79.05 tokens
- max: 207 tokens
- min: 28 tokens
- mean: 80.26 tokens
- max: 256 tokens
- min: 21 tokens
- mean: 63.02 tokens
- max: 184 tokens
- Samples:
sentence positive_sentence negative_sentence भारतका रक्षा प्रमुख जनरल अनिल चौहानले चीन, पाकिस्तान र बंगलादेशको गठबन्धनलाई भारतको सुरक्षाका लागि ठूलो खतरा भएको बताएका छन्।चीन, पाकिस्तान र बंगलादेशको संयुक्त गठबन्धनलाई भारतको सुरक्षा चुनौतीको रूपमा जनरल अनिल चौहानले औंल्याएका छन्।जनरल अनिल चौहानले भारतको सुरक्षाका लागि चीन र पाकिस्तानको गठबन्धन भन्दा मात्र बंगलादेशलाई ठूलो खतरा भनेका छन्।The Sarlahi District Court has granted bail to suspended Bagmati Municipality Mayor Bharat Bahadur Thapa on a Rs 5 million bond in connection with illegal forest resource extraction charges.Bharat Bahadur Thapa, the suspended mayor of Bagmati Municipality, was released on bail by the Sarlahi District Court after posting a bond of Rs 5 million related to allegations of unlawful forest resource exploitation.The Sarlahi District Court denied bail to Bharat Bahadur Thapa, the suspended mayor of Bagmati Municipality, in the case concerning illegal forest resource extraction.नेपाली युवा महिला फुटबल टोलीका प्रशिक्षक यामप्रसाद गुरुङले साफ यु-२० महिला च्याम्पियनसिपको उपाधि जित्ने आशा दिएका छन्। नेपालले जुलाईमा बंगलादेशमा हुने प्रतियोगितामा बलियो टोली बनाएर प्रतिस्पर्धा गर्नेछ।यामप्रसाद गुरुङले साफ यु-२० महिला च्याम्पियनसिपमा नेपालको टोलीले शीर्ष स्थान हासिल गर्ने विश्वास व्यक्त गरेका छन्। आगामी जुलाईमा बंगलादेशमा आयोजना हुने प्रतियोगितामा नेपालले सशक्त टोली प्रस्तुत गर्ने तयारीमा छ।नेपाली युवा महिला फुटबल टोलीले आगामी साफ यु-२० महिला च्याम्पियनसिपमा कमजोर प्रदर्शन गर्ने सम्भावना व्यक्त गरिएको छ। - Loss:
TripletLosswith these parameters:{ "distance_metric": "TripletDistanceMetric.COSINE", "triplet_margin": 0.2 }
stsb_en
stsb_en
- Dataset: stsb_en
- Size: 5,710 training samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 6 tokens
- mean: 10.03 tokens
- max: 28 tokens
- min: 5 tokens
- mean: 9.96 tokens
- max: 25 tokens
- min: 0.0
- mean: 0.45
- max: 1.0
- Samples:
sentence1 sentence2 score A plane is taking off.An air plane is taking off.1.0A man is playing a large flute.A man is playing a flute.0.76A man is spreading shreded cheese on a pizza.A man is spreading shredded cheese on an uncooked pizza.0.76 - Loss:
CosineSimilarityLosswith these parameters:{ "loss_fct": "torch.nn.modules.loss.MSELoss" }
stsb_ne
stsb_ne
- Dataset: stsb_ne
- Size: 5,710 training samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 6 tokens
- mean: 21.82 tokens
- max: 60 tokens
- min: 6 tokens
- mean: 21.96 tokens
- max: 54 tokens
- min: 0.0
- mean: 0.45
- max: 1.0
- Samples:
sentence1 sentence2 score एउटा विमान उडिरहेको छ।हवाई जहाज उडिरहेको छ।1.0एउटा मान्छे ठूलो बाँसुरी बजाइरहेको छ।एउटा मान्छे बाँसुरी बजाउँदै छ।0.76एक मानिस पिज्जामा टुक्रा चिज फैलाउँदै छ।एक जना मानिसले न पकाएको पिज्जामा टुक्रा पारेको चीज फैलाउँदै छ।0.76 - Loss:
CoSENTLosswith these parameters:{ "scale": 20.0, "similarity_fct": "pairwise_cos_sim" }
stsb_en_ne
stsb_en_ne
- Dataset: stsb_en_ne
- Size: 5,710 training samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 6 tokens
- mean: 10.03 tokens
- max: 28 tokens
- min: 6 tokens
- mean: 21.96 tokens
- max: 54 tokens
- min: 0.0
- mean: 0.45
- max: 1.0
- Samples:
sentence1 sentence2 score A plane is taking off.हवाई जहाज उडिरहेको छ।1.0A man is playing a large flute.एउटा मान्छे बाँसुरी बजाउँदै छ।0.76A man is spreading shreded cheese on a pizza.एक जना मानिसले न पकाएको पिज्जामा टुक्रा पारेको चीज फैलाउँदै छ।0.76 - Loss:
CoSENTLosswith these parameters:{ "scale": 20.0, "similarity_fct": "pairwise_cos_sim" }
stsb_ne_en
stsb_ne_en
- Dataset: stsb_ne_en
- Size: 5,710 training samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 6 tokens
- mean: 21.82 tokens
- max: 60 tokens
- min: 5 tokens
- mean: 9.96 tokens
- max: 25 tokens
- min: 0.0
- mean: 0.45
- max: 1.0
- Samples:
sentence1 sentence2 score एउटा विमान उडिरहेको छ।An air plane is taking off.1.0एउटा मान्छे ठूलो बाँसुरी बजाइरहेको छ।A man is playing a flute.0.76एक मानिस पिज्जामा टुक्रा चिज फैलाउँदै छ।A man is spreading shredded cheese on an uncooked pizza.0.76 - Loss:
CoSENTLosswith these parameters:{ "scale": 20.0, "similarity_fct": "pairwise_cos_sim" }
Evaluation Datasets
title_excerpt
title_excerpt
- Dataset: title_excerpt
- Size: 4,333 evaluation samples
- Columns:
titleandexcerpt - Approximate statistics based on the first 1000 samples:
title excerpt type string string details - min: 6 tokens
- mean: 28.18 tokens
- max: 77 tokens
- min: 24 tokens
- mean: 73.83 tokens
- max: 183 tokens
- Samples:
title excerpt बडीमालिकाका सात वडा छाउपडी गोठमुक्तबडीमालिका नगरपालिकाका सात वटा वडाहरु छाउपडी प्रथाबाट मुक्त घोषणा गरिएका छन्। ९० प्रतिशत घर छाउ गोठमुक्त भएपछि वडा नम्बर ८ लाई नयाँ छाउपडी मुक्त वडा घोषणा गरिएको छ।What are Iran’s ballistic missile capabilities?Iran possesses one of the largest ballistic missile arsenals in the Middle East, with missiles capable of reaching Israel and beyond. Recent developments include the introduction of hypersonic missiles and underground missile facilities, underscoring Iran's strategic deterrence ambitions.पर्सामा ट्रयाक्टर दुर्घटनामा परेर दुईको मृत्यु, एक घाइतेपर्सामा ट्रयाक्टर दुर्घटनामा परेर दुई जनाको मृत्यु भएको छ भने एक जना घाइते भएका छन्। - Loss:
MultipleNegativesSymmetricRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim" }
ne_en
ne_en
- Dataset: ne_en
- Size: 189 evaluation samples
- Columns:
titleandtranslation - Approximate statistics based on the first 189 samples:
title translation type string string details - min: 7 tokens
- mean: 29.78 tokens
- max: 71 tokens
- min: 5 tokens
- mean: 23.54 tokens
- max: 85 tokens
- Samples:
title translation House passes Federal Civil Service Billसंघीय निजामती सेवा विधेयक संसदबाट पारितघूस लेनदेन : कांग्रेसबाट निर्वाचित नागार्जुनका मेयर बस्नेतलाई ८ वर्ष जेल र २ करोड ३० लाख जरिवानाBribery transaction: Mayor Basnet of Nagarjuna, who was elected by the Congress, was sentenced to 8 years in prison and a fine of 23 millionकुलिङ पिरियडको कमजोरीको नैतिक जिम्मेवारी लिन्छु, दोषीमाथि कारवाही हुनुपर्छ : रामहरि खतिवडाI take moral responsibility for the weakness of the cooling period, action should be taken against the guilty: Ramhari Khatiwada - Loss:
MultipleNegativesSymmetricRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim" }
excerpt_paraphrase
excerpt_paraphrase
- Dataset: excerpt_paraphrase
- Size: 27 evaluation samples
- Columns:
sentence1andsentence2 - Approximate statistics based on the first 27 samples:
sentence1 sentence2 type string string details - min: 33 tokens
- mean: 83.3 tokens
- max: 141 tokens
- min: 35 tokens
- mean: 90.89 tokens
- max: 167 tokens
- Samples:
sentence1 sentence2 त्रिभुवन विश्वविद्यालय परीक्षा नियन्त्रण कार्यालयले दश वटा मुख्य सेवा अनलाइनमा ल्याएपछि विद्यार्थीले कार्यालय धाउनु नपर्ने व्यवस्था भएको छ। अब विद्यार्थीले घरबाटै अनलाइन आवेदन दिई प्रमाणपत्र प्राप्त गर्न सक्नेछन्।त्रिभुवन विश्वविद्यालयको परीक्षा नियन्त्रण कार्यालयले दश प्रमुख सेवाहरू अनलाइनमा सारेपछि विद्यार्थीहरूले कार्यालय धाउनु पर्ने झन्झटबाट मुक्ति पाएका छन्। यसले गर्दा अब उनीहरूले घरबाटै अनलाइन आवेदन दिएर प्रमाणपत्रहरू प्राप्त गर्न सक्नेछन्।नेपाल मेडिसिटी अस्पताल ललितपुरमा पहिलो पटक कलेजो प्रत्यारोपण सफल रुपमा गरिएको छ, जहाँ ईश्वर कार्कीलाई उनका छोरा आयुष कार्कीले कलेजो दान गरेका छन्।ललितपुरस्थित नेपाल मेडिसिटी अस्पतालमा पहिलो सफल कलेजो प्रत्यारोपण सम्पन्न भएको छ, जसमा आयुष कार्कीले आफ्ना बुबा ईश्वर कार्कीलाई कलेजो प्रदान गरेका थिए।अमेरिकी राष्ट्रपति डोनाल्ड ट्रम्पले इजरायल र इरान बीच युद्धविराम कार्यान्वयनमा आएको बताएका छन् र दुवै देशलाई उल्लंघन नगर्न आग्रह गरेका छन्।अमेरिकी राष्ट्रपति डोनाल्ड ट्रम्पले इजरायल र इरानबीच युद्धविराम लागू भएको घोषणा गर्दै दुवै पक्षलाई त्यसको पालना गर्न आग्रह गरेका छन्। - Loss:
MultipleNegativesSymmetricRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim" }
nepali_triplets
nepali_triplets
- Dataset: nepali_triplets
- Size: 200 evaluation samples
- Columns:
sentence,positive_sentence, andnegative_sentence - Approximate statistics based on the first 200 samples:
sentence positive_sentence negative_sentence type string string string details - min: 25 tokens
- mean: 76.12 tokens
- max: 177 tokens
- min: 27 tokens
- mean: 77.06 tokens
- max: 151 tokens
- min: 21 tokens
- mean: 61.07 tokens
- max: 120 tokens
- Samples:
sentence positive_sentence negative_sentence नेपालको सेयर बजारमा मंगलबार ४ अंकले गिरावट आएको छ र कारोबार रकम १३ अर्ब १८ करोडबाट ७ अर्ब ६५ करोडमा घटेको छ। पुरे इनर्जी लिमिटेडको सेयरमा सर्किट लागेको छ।मंगलबार नेपाल स्टक मार्केटमा ४ अंकको कमी देखिएको छ भने कारोबार रकम १३ अर्ब १८ करोडबाट घटेर ७ अर्ब ६५ करोड पुगेको छ। पुरे इनर्जी लिमिटेडको सेयरमा सर्किट ब्रेकर लाग्न पुगेको छ।नेपालको सेयर बजारमा मंगलबार ४ अंकले वृद्धि भएको छ र कारोबार रकम ७ अर्ब ६५ करोडबाट १३ अर्ब १८ करोडमा बढेको छ। पुरे इनर्जी लिमिटेडको सेयरमा कुनै सर्किट लागेको छैन।नेपाल चेम्बर अफ कमर्सले सुन र बहुमूल्य पत्थरमा लगाइएको विलासिता कर र भ्याट कर पुनरवलोकन गर्न अर्थमन्त्री विष्णु प्रसाद पौडेललाई अनुरोध गर्यो।अर्थमन्त्री विष्णु प्रसाद पौडेललाई नेपाल चेम्बर अफ कमर्सले सुन र बहुमूल्य पत्थरमा लाग्ने विलासिता कर र भ्याट करको समीक्षा गर्न आग्रह गरेको छ।नेपाल चेम्बर अफ कमर्सले सुन र बहुमूल्य पत्थरमा लगाइएको करहरू बढाउन अर्थमन्त्री विष्णु प्रसाद पौडेलसँग अनुरोध गरेको छ।The CIAA arrested the chief and an assistant of the Bhaktapur Land Revenue Office for accepting a bribe of Rs 2.45 million.Authorities from the CIAA took into custody the head and a subordinate of Bhaktapur's Land Revenue Office on charges of receiving a Rs 2.45 million bribe.The CIAA investigated the Bhaktapur Land Revenue Office chief and assistant for alleged negligence in land record management. - Loss:
TripletLosswith these parameters:{ "distance_metric": "TripletDistanceMetric.COSINE", "triplet_margin": 0.2 }
stsb_en
stsb_en
- Dataset: stsb_en
- Size: 1,378 evaluation samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 5 tokens
- mean: 13.51 tokens
- max: 43 tokens
- min: 5 tokens
- mean: 13.47 tokens
- max: 42 tokens
- min: 0.0
- mean: 0.5
- max: 1.0
- Samples:
sentence1 sentence2 score A girl is styling her hair.A girl is brushing her hair.0.5A group of men play soccer on the beach.A group of boys are playing soccer on the beach.0.72One woman is measuring another woman's ankle.A woman measures another woman's ankle.1.0 - Loss:
CosineSimilarityLosswith these parameters:{ "loss_fct": "torch.nn.modules.loss.MSELoss" }
stsb_ne
stsb_ne
- Dataset: stsb_ne
- Size: 1,378 evaluation samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 7 tokens
- mean: 29.79 tokens
- max: 147 tokens
- min: 7 tokens
- mean: 30.01 tokens
- max: 127 tokens
- min: 0.0
- mean: 0.5
- max: 1.0
- Samples:
sentence1 sentence2 score एउटी केटी आफ्नो कपाल स्टाइल गर्दै छ।एउटी केटी आफ्नो कपाल माझ्दै छ।0.5पुरुषहरूको समूह समुद्र तटमा फुटबल खेल्दै छ।केटाहरूको समूह समुद्र तटमा फुटबल खेल्दै छ।0.72एउटी महिलाले अर्की महिलाको घुँडा नाप्दै छिन्।एउटी महिलाले अर्को महिलाको खुट्टा नाप्छिन्।1.0 - Loss:
CoSENTLosswith these parameters:{ "scale": 20.0, "similarity_fct": "pairwise_cos_sim" }
stsb_en_ne
stsb_en_ne
- Dataset: stsb_en_ne
- Size: 1,378 evaluation samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 5 tokens
- mean: 13.51 tokens
- max: 43 tokens
- min: 7 tokens
- mean: 30.01 tokens
- max: 127 tokens
- min: 0.0
- mean: 0.5
- max: 1.0
- Samples:
sentence1 sentence2 score A girl is styling her hair.एउटी केटी आफ्नो कपाल माझ्दै छ।0.5A group of men play soccer on the beach.केटाहरूको समूह समुद्र तटमा फुटबल खेल्दै छ।0.72One woman is measuring another woman's ankle.एउटी महिलाले अर्को महिलाको खुट्टा नाप्छिन्।1.0 - Loss:
CoSENTLosswith these parameters:{ "scale": 20.0, "similarity_fct": "pairwise_cos_sim" }
stsb_ne_en
stsb_ne_en
- Dataset: stsb_ne_en
- Size: 1,378 evaluation samples
- Columns:
sentence1,sentence2, andscore - Approximate statistics based on the first 1000 samples:
sentence1 sentence2 score type string string float details - min: 7 tokens
- mean: 29.79 tokens
- max: 147 tokens
- min: 5 tokens
- mean: 13.47 tokens
- max: 42 tokens
- min: 0.0
- mean: 0.5
- max: 1.0
- Samples:
sentence1 sentence2 score एउटी केटी आफ्नो कपाल स्टाइल गर्दै छ।A girl is brushing her hair.0.5पुरुषहरूको समूह समुद्र तटमा फुटबल खेल्दै छ।A group of boys are playing soccer on the beach.0.72एउटी महिलाले अर्की महिलाको घुँडा नाप्दै छिन्।A woman measures another woman's ankle.1.0 - Loss:
CoSENTLosswith these parameters:{ "scale": 20.0, "similarity_fct": "pairwise_cos_sim" }
Training Hyperparameters
Non-Default Hyperparameters
eval_strategy: stepsper_device_train_batch_size: 64learning_rate: 2e-05num_train_epochs: 1warmup_ratio: 0.1load_best_model_at_end: Truegradient_checkpointing: True
All Hyperparameters
Click to expand
overwrite_output_dir: Falsedo_predict: Falseeval_strategy: stepsprediction_loss_only: Trueper_device_train_batch_size: 64per_device_eval_batch_size: 8per_gpu_train_batch_size: Noneper_gpu_eval_batch_size: Nonegradient_accumulation_steps: 1eval_accumulation_steps: Nonetorch_empty_cache_steps: Nonelearning_rate: 2e-05weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0num_train_epochs: 1max_steps: -1lr_scheduler_type: linearlr_scheduler_kwargs: {}warmup_ratio: 0.1warmup_steps: 0log_level: passivelog_level_replica: warninglog_on_each_node: Truelogging_nan_inf_filter: Truesave_safetensors: Truesave_on_each_node: Falsesave_only_model: Falserestore_callback_states_from_checkpoint: Falseno_cuda: Falseuse_cpu: Falseuse_mps_device: Falseseed: 42data_seed: Nonejit_mode_eval: Falsebf16: Falsefp16: Falsefp16_opt_level: O1half_precision_backend: autobf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonelocal_rank: 0ddp_backend: Nonetpu_num_cores: Nonetpu_metrics_debug: Falsedebug: []dataloader_drop_last: Falsedataloader_num_workers: 0dataloader_prefetch_factor: Nonepast_index: -1disable_tqdm: Falseremove_unused_columns: Truelabel_names: Noneload_best_model_at_end: Trueignore_data_skip: Falsefsdp: []fsdp_min_num_params: 0fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}fsdp_transformer_layer_cls_to_wrap: Noneaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedeepspeed: Nonelabel_smoothing_factor: 0.0optim: adamw_torch_fusedoptim_args: Noneadafactor: Falsegroup_by_length: Falselength_column_name: lengthproject: huggingfacetrackio_space_id: trackioddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falsedataloader_pin_memory: Truedataloader_persistent_workers: Falseskip_memory_metrics: Trueuse_legacy_prediction_loop: Falsepush_to_hub: Falseresume_from_checkpoint: Nonehub_model_id: Nonehub_strategy: every_savehub_private_repo: Nonehub_always_push: Falsehub_revision: Nonegradient_checkpointing: Truegradient_checkpointing_kwargs: Noneinclude_inputs_for_metrics: Falseinclude_for_metrics: []eval_do_concat_batches: Truefp16_backend: autopush_to_hub_model_id: Nonepush_to_hub_organization: Nonemp_parameters:auto_find_batch_size: Falsefull_determinism: Falsetorchdynamo: Noneray_scope: lastddp_timeout: 1800torch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneinclude_tokens_per_second: Falseinclude_num_input_tokens_seen: noneftune_noise_alpha: Noneoptim_target_modules: Nonebatch_eval_metrics: Falseeval_on_start: Falseuse_liger_kernel: Falseliger_kernel_config: Noneeval_use_gather_object: Falseaverage_tokens_across_devices: Trueprompts: Nonebatch_sampler: batch_samplermulti_dataset_batch_sampler: proportional
Training Logs
| Epoch | Step | Training Loss | title excerpt loss | ne en loss | excerpt paraphrase loss | nepali triplets loss | stsb en loss | stsb ne loss | stsb en ne loss | stsb ne en loss | multi_lang_ir_cosine_ndcg@10 | en_ir_cosine_ndcg@10 | ne_ir_cosine_ndcg@10 | translation_mean_accuracy | nepali_triplets_cosine_accuracy | stsb_en_spearman_cosine | stsb_ne_spearman_cosine |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.0553 | 100 | 3.1763 | 0.0438 | 0.2323 | 0.0001 | 0.1189 | 0.0389 | 6.6070 | 6.1735 | 6.2053 | 0.7926 | 0.8915 | 0.7251 | 0.4947 | 0.4750 | 0.8671 | 0.6586 |
| 0.1107 | 200 | 1.8472 | 0.0338 | 0.2238 | 0.0002 | 0.1076 | 0.0376 | 6.3000 | 5.8402 | 5.8531 | 0.8124 | 0.8999 | 0.7522 | 0.5020 | 0.5250 | 0.8700 | 0.6681 |
| 0.1660 | 300 | 2.1218 | 0.0285 | 0.2096 | 0.0002 | 0.1007 | 0.0362 | 6.3730 | 5.8704 | 5.7911 | 0.8131 | 0.9043 | 0.7497 | 0.5073 | 0.5450 | 0.8739 | 0.6651 |
| 0.2214 | 400 | 1.4009 | 0.0262 | 0.2184 | 0.0001 | 0.1035 | 0.0361 | 6.4156 | 5.8561 | 5.7325 | 0.8204 | 0.9082 | 0.7593 | 0.5053 | 0.5350 | 0.8758 | 0.6630 |
| 0.2767 | 500 | 2.4768 | 0.0243 | 0.2268 | 0.0001 | 0.1030 | 0.0362 | 6.2723 | 5.6653 | 5.5011 | 0.8255 | 0.9091 | 0.7669 | 0.4814 | 0.5100 | 0.8749 | 0.6685 |
| 0.3320 | 600 | 1.6207 | 0.0247 | 0.1964 | 0.0001 | 0.1078 | 0.0348 | 6.2414 | 5.5925 | 5.4461 | 0.8282 | 0.9106 | 0.7708 | 0.5 | 0.5 | 0.8765 | 0.6709 |
| 0.3874 | 700 | 2.317 | 0.0246 | 0.1865 | 0.0001 | 0.1012 | 0.0353 | 6.2366 | 5.5786 | 5.3861 | 0.8316 | 0.9106 | 0.7764 | 0.5166 | 0.5400 | 0.8758 | 0.6718 |
| 0.4427 | 800 | 1.0611 | 0.0236 | 0.2016 | 0.0001 | 0.0984 | 0.0344 | 6.3357 | 5.6502 | 5.4199 | 0.8330 | 0.9121 | 0.7775 | 0.5020 | 0.5600 | 0.8787 | 0.6732 |
| 0.4981 | 900 | 1.4653 | 0.0221 | 0.1887 | 0.0001 | 0.0990 | 0.0334 | 6.3109 | 5.5926 | 5.4254 | 0.8335 | 0.9130 | 0.7778 | 0.5153 | 0.5300 | 0.8788 | 0.6737 |
| 0.5534 | 1000 | 1.3721 | 0.0228 | 0.1501 | 0.0001 | 0.1021 | 0.0342 | 6.3623 | 5.5621 | 5.3895 | 0.8337 | 0.9145 | 0.7773 | 0.5505 | 0.5350 | 0.8794 | 0.6701 |
| 0.6087 | 1100 | 2.0056 | 0.0212 | 0.1564 | 0.0001 | 0.1022 | 0.0344 | 6.2683 | 5.5163 | 5.3413 | 0.8367 | 0.9147 | 0.7820 | 0.5332 | 0.5250 | 0.8790 | 0.6740 |
| 0.6641 | 1200 | 1.4577 | 0.0212 | 0.1500 | 0.0001 | 0.1002 | 0.0342 | 6.2664 | 5.4797 | 5.3139 | 0.8364 | 0.9146 | 0.7817 | 0.5418 | 0.5250 | 0.8787 | 0.6742 |
| 0.7194 | 1300 | 1.8814 | 0.0208 | 0.1645 | 0.0001 | 0.1013 | 0.0339 | 6.2017 | 5.4290 | 5.2550 | 0.8389 | 0.9154 | 0.7852 | 0.5232 | 0.5250 | 0.8794 | 0.6762 |
| 0.7748 | 1400 | 1.9753 | 0.0206 | 0.1562 | 0.0001 | 0.0997 | 0.0339 | 6.1869 | 5.3668 | 5.2240 | 0.8396 | 0.9153 | 0.7864 | 0.5398 | 0.5300 | 0.8794 | 0.6783 |
| 0.8301 | 1500 | 1.3772 | 0.0206 | 0.1478 | 0.0001 | 0.1022 | 0.0339 | 6.2240 | 5.3975 | 5.2144 | 0.8397 | 0.9155 | 0.7866 | 0.5471 | 0.5150 | 0.8792 | 0.6770 |
| 0.8854 | 1600 | 2.3584 | 0.0206 | 0.1465 | 0.0001 | 0.1009 | 0.0344 | 6.1776 | 5.3378 | 5.1625 | 0.8412 | 0.9162 | 0.7888 | 0.5445 | 0.5250 | 0.8791 | 0.6785 |
| 0.9408 | 1700 | 2.5933 | 0.0206 | 0.1444 | 0.0001 | 0.1001 | 0.0345 | 6.1674 | 5.3130 | 5.1515 | 0.8413 | 0.9160 | 0.7890 | 0.5511 | 0.5350 | 0.8791 | 0.6788 |
| 0.9961 | 1800 | 1.6188 | 0.0207 | 0.1420 | 0.0001 | 0.0996 | 0.0344 | 6.1701 | 5.3158 | 5.1587 | 0.8412 | 0.9163 | 0.7887 | 0.5571 | 0.5400 | 0.8792 | 0.6788 |
| -1 | -1 | - | - | - | - | - | - | - | - | - | 0.8412 | 0.9163 | 0.7886 | 0.5571 | 0.5400 | 0.8791 | 0.6788 |
Framework Versions
- Python: 3.11.13
- Sentence Transformers: 4.1.0
- Transformers: 4.57.2
- PyTorch: 2.12.0
- Accelerate: 1.13.0
- Datasets: 4.8.5
- Tokenizers: 0.22.2
Citation
BibTeX
If you use this model, please cite it as:
@misc{subedi2026allminilml6v3nepali,
author = {Subedi, Sanjaya},
title = {all-MiniLM-L6-v3-nepali: A Nepali Sentence Transformer for Semantic Search and Similarity},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face model repository},
howpublished = {\url{https://huggingface.co/jangedoo/all-MiniLM-L6-v3-nepali}},
note = {Fine-tuned Sentence Transformer model for Nepali and English-Nepali semantic similarity, semantic search, paraphrase mining, clustering, and retrieval}
}
- Downloads last month
- 79
Model tree for jangedoo/all-MiniLM-L6-v3-nepali
Unable to build the model tree, the base model loops to the model itself. Learn more.
Evaluation results
- Cosine Accuracy@10 on multi lang irself-reported0.917
- Cosine Precision@10 on multi lang irself-reported0.092
- Cosine Precision@50 on multi lang irself-reported0.019
- Cosine Recall@10 on multi lang irself-reported0.917
- Cosine Recall@50 on multi lang irself-reported0.967
- Cosine Ndcg@10 on multi lang irself-reported0.841
- Cosine Mrr@10 on multi lang irself-reported0.816
- Cosine Map@100 on multi lang irself-reported0.819