How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("feature-extraction", model="ai-forever/FRIDA")
# Load model directly
from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("ai-forever/FRIDA")
model = AutoModel.from_pretrained("ai-forever/FRIDA", device_map="auto")
Quick Links

Model Card for FRIDA

FRIDA is a full-scale finetuned general text embedding model inspired by denoising architecture based on T5. The model is based on the encoder part of FRED-T5 model and continues research of text embedding models (ruMTEB, ru-en-RoSBERTa). It has been pre-trained on a Russian-English dataset and fine-tuned for improved performance on the target task.

For more model details please refer to our article (RU). The model's results are presented on the MTEB and rusBEIR leaderboards.

Usage

The model can be used as is with prefixes. It is recommended to use CLS pooling. The choice of prefix and pooling depends on the task.

We use the following basic rules to choose a prefix:

  • "search_query: " and "search_document: " prefixes are for answer or relevant paragraph retrieval
  • "paraphrase: " prefix is for symmetric paraphrasing related tasks (STS, paraphrase mining, deduplication)
  • "categorize: " prefix is for asymmetric matching of document title and body (e.g. news, scientific papers, social posts)
  • "categorize_sentiment: " prefix is for any tasks that rely on sentiment features (e.g. hate, toxic, emotion)
  • "categorize_topic: " prefix is intended for tasks where you need to group texts by topic
  • "categorize_entailment: " prefix is for textual entailment task (NLI)

To better tailor the model to your needs, you can fine-tune it with relevant high-quality Russian and English datasets.

Below are examples of texts encoding using the Transformers and SentenceTransformers libraries.

Transformers

import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, T5EncoderModel


def pool(hidden_state, mask, pooling_method="cls"):
    if pooling_method == "mean":
        s = torch.sum(hidden_state * mask.unsqueeze(-1).float(), dim=1)
        d = mask.sum(axis=1, keepdim=True).float()
        return s / d
    elif pooling_method == "cls":
        return hidden_state[:, 0]

inputs = [
    # 
    "paraphrase: Π’ Ярославской области Ρ€Π°Π·Ρ€Π΅ΡˆΠΈΠ»ΠΈ Ρ€Π°Π±ΠΎΡ‚Ρƒ бань, Π½ΠΎ Π±Π΅Π· посСтитСлСй",
    "categorize_entailment: Π–Π΅Π½Ρ‰ΠΈΠ½Ρƒ доставили Π² Π±ΠΎΠ»ΡŒΠ½ΠΈΡ†Ρƒ, Π·Π° Π΅Π΅ Тизнь сСйчас Π±ΠΎΡ€ΡŽΡ‚ΡΡ Π²Ρ€Π°Ρ‡ΠΈ.",
    "search_query: Бколько программистов Π½ΡƒΠΆΠ½ΠΎ, Ρ‡Ρ‚ΠΎΠ±Ρ‹ Π²ΠΊΡ€ΡƒΡ‚ΠΈΡ‚ΡŒ Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΡƒ?",
    # 
    "paraphrase: Ярославским баням Ρ€Π°Π·Ρ€Π΅ΡˆΠΈΠ»ΠΈ Ρ€Π°Π±ΠΎΡ‚Π°Ρ‚ΡŒ Π±Π΅Π· посСтитСлСй",
    "categorize_entailment: Π–Π΅Π½Ρ‰ΠΈΠ½Ρƒ ΡΠΏΠ°ΡΠ°ΡŽΡ‚ Π²Ρ€Π°Ρ‡ΠΈ.",
    "search_document: Π§Ρ‚ΠΎΠ±Ρ‹ Π²ΠΊΡ€ΡƒΡ‚ΠΈΡ‚ΡŒ Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΡƒ, трСбуСтся Ρ‚Ρ€ΠΈ программиста: ΠΎΠ΄ΠΈΠ½ Π½Π°ΠΏΠΈΡˆΠ΅Ρ‚ ΠΏΡ€ΠΎΠ³Ρ€Π°ΠΌΠΌΡƒ извлСчСния Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΠΈ, Π΄Ρ€ΡƒΠ³ΠΎΠΉ β€” вкручивания Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΠΈ, Π° Ρ‚Ρ€Π΅Ρ‚ΠΈΠΉ ΠΏΡ€ΠΎΠ²Π΅Π΄Π΅Ρ‚ тСстированиС."
]

tokenizer = AutoTokenizer.from_pretrained("ai-forever/FRIDA")
model = T5EncoderModel.from_pretrained("ai-forever/FRIDA")

tokenized_inputs = tokenizer(inputs, max_length=512, padding=True, truncation=True, return_tensors="pt")

with torch.no_grad():
    outputs = model(**tokenized_inputs)
    
embeddings = pool(
    outputs.last_hidden_state, 
    tokenized_inputs["attention_mask"],
    pooling_method="cls" # or try "mean"
)

embeddings = F.normalize(embeddings, p=2, dim=1)
sim_scores = embeddings[:3] @ embeddings[3:].T
print(sim_scores.diag().tolist())
# [0.9360030293464661, 0.8591322302818298, 0.728583037853241]

SentenceTransformers

from sentence_transformers import SentenceTransformer

inputs = [
    # 
    "paraphrase: Π’ Ярославской области Ρ€Π°Π·Ρ€Π΅ΡˆΠΈΠ»ΠΈ Ρ€Π°Π±ΠΎΡ‚Ρƒ бань, Π½ΠΎ Π±Π΅Π· посСтитСлСй",
    "categorize_entailment: Π–Π΅Π½Ρ‰ΠΈΠ½Ρƒ доставили Π² Π±ΠΎΠ»ΡŒΠ½ΠΈΡ†Ρƒ, Π·Π° Π΅Π΅ Тизнь сСйчас Π±ΠΎΡ€ΡŽΡ‚ΡΡ Π²Ρ€Π°Ρ‡ΠΈ.",
    "search_query: Бколько программистов Π½ΡƒΠΆΠ½ΠΎ, Ρ‡Ρ‚ΠΎΠ±Ρ‹ Π²ΠΊΡ€ΡƒΡ‚ΠΈΡ‚ΡŒ Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΡƒ?",
    # 
    "paraphrase: Ярославским баням Ρ€Π°Π·Ρ€Π΅ΡˆΠΈΠ»ΠΈ Ρ€Π°Π±ΠΎΡ‚Π°Ρ‚ΡŒ Π±Π΅Π· посСтитСлСй",
    "categorize_entailment: Π–Π΅Π½Ρ‰ΠΈΠ½Ρƒ ΡΠΏΠ°ΡΠ°ΡŽΡ‚ Π²Ρ€Π°Ρ‡ΠΈ.",
    "search_document: Π§Ρ‚ΠΎΠ±Ρ‹ Π²ΠΊΡ€ΡƒΡ‚ΠΈΡ‚ΡŒ Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΡƒ, трСбуСтся Ρ‚Ρ€ΠΈ программиста: ΠΎΠ΄ΠΈΠ½ Π½Π°ΠΏΠΈΡˆΠ΅Ρ‚ ΠΏΡ€ΠΎΠ³Ρ€Π°ΠΌΠΌΡƒ извлСчСния Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΠΈ, Π΄Ρ€ΡƒΠ³ΠΎΠΉ β€” вкручивания Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΠΈ, Π° Ρ‚Ρ€Π΅Ρ‚ΠΈΠΉ ΠΏΡ€ΠΎΠ²Π΅Π΄Π΅Ρ‚ тСстированиС."
]

# loads model with CLS pooling
model = SentenceTransformer("ai-forever/FRIDA")

# embeddings are normalized by default
embeddings = model.encode(inputs, convert_to_tensor=True)

sim_scores = embeddings[:3] @ embeddings[3:].T
print(sim_scores.diag().tolist())
# [0.9360026717185974, 0.8591331243515015, 0.7285830974578857]

or using prompts (sentence-transformers>=2.4.0):

from sentence_transformers import SentenceTransformer

# loads model with CLS pooling
model = SentenceTransformer("ai-forever/FRIDA")

paraphrase = model.encode(["Π’ Ярославской области Ρ€Π°Π·Ρ€Π΅ΡˆΠΈΠ»ΠΈ Ρ€Π°Π±ΠΎΡ‚Ρƒ бань, Π½ΠΎ Π±Π΅Π· посСтитСлСй", "Ярославским баням Ρ€Π°Π·Ρ€Π΅ΡˆΠΈΠ»ΠΈ Ρ€Π°Π±ΠΎΡ‚Π°Ρ‚ΡŒ Π±Π΅Π· посСтитСлСй"], prompt_name="paraphrase")
print(paraphrase[0] @ paraphrase[1].T) # 0.9360032

categorize_entailment = model.encode(["Π–Π΅Π½Ρ‰ΠΈΠ½Ρƒ доставили Π² Π±ΠΎΠ»ΡŒΠ½ΠΈΡ†Ρƒ, Π·Π° Π΅Π΅ Тизнь сСйчас Π±ΠΎΡ€ΡŽΡ‚ΡΡ Π²Ρ€Π°Ρ‡ΠΈ.", "Π–Π΅Π½Ρ‰ΠΈΠ½Ρƒ ΡΠΏΠ°ΡΠ°ΡŽΡ‚ Π²Ρ€Π°Ρ‡ΠΈ."], prompt_name="categorize_entailment")
print(categorize_entailment[0] @ categorize_entailment[1].T) # 0.8591322

query_embedding = model.encode("Бколько программистов Π½ΡƒΠΆΠ½ΠΎ, Ρ‡Ρ‚ΠΎΠ±Ρ‹ Π²ΠΊΡ€ΡƒΡ‚ΠΈΡ‚ΡŒ Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΡƒ?", prompt_name="search_query")
document_embedding = model.encode("Π§Ρ‚ΠΎΠ±Ρ‹ Π²ΠΊΡ€ΡƒΡ‚ΠΈΡ‚ΡŒ Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΡƒ, трСбуСтся Ρ‚Ρ€ΠΈ программиста: ΠΎΠ΄ΠΈΠ½ Π½Π°ΠΏΠΈΡˆΠ΅Ρ‚ ΠΏΡ€ΠΎΠ³Ρ€Π°ΠΌΠΌΡƒ извлСчСния Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΠΈ, Π΄Ρ€ΡƒΠ³ΠΎΠΉ β€” вкручивания Π»Π°ΠΌΠΏΠΎΡ‡ΠΊΠΈ, Π° Ρ‚Ρ€Π΅Ρ‚ΠΈΠΉ ΠΏΡ€ΠΎΠ²Π΅Π΄Π΅Ρ‚ тСстированиС.", prompt_name="search_document")
print(query_embedding @ document_embedding.T) # 0.7285831

Results

FRIDA is the top-1 model among models with up to 3 billion parameters (11.08.26).

Authors

Citation

@misc{TODO
}

Limitations

The model is designed to process texts in Russian, the quality in English is unknown. Maximum input text length is limited to 512 tokens.

Downloads last month
123,385
Safetensors
Model size
0.8B params
Tensor type
F32
Β·
Inference Providers NEW

Model tree for ai-forever/FRIDA

Finetuned
(7)
this model
Finetunes
1 model
Quantizations
2 models

Dataset used to train ai-forever/FRIDA

Spaces using ai-forever/FRIDA 18

Collection including ai-forever/FRIDA

Papers for ai-forever/FRIDA