Instructions to use Kansallisarkisto/finbert-ner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kansallisarkisto/finbert-ner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="Kansallisarkisto/finbert-ner")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("Kansallisarkisto/finbert-ner") model = AutoModelForTokenClassification.from_pretrained("Kansallisarkisto/finbert-ner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update example code in README
Browse files
README.md
CHANGED
|
@@ -52,16 +52,25 @@ contains for example old names or writing styles.
|
|
| 52 |
The easiest way to use the model is by utilizing the Transformers pipeline for token classification:
|
| 53 |
|
| 54 |
```python
|
| 55 |
-
from transformers import pipeline
|
| 56 |
|
| 57 |
model_checkpoint = "Kansallisarkisto/finbert-ner"
|
|
|
|
|
|
|
|
|
|
| 58 |
token_classifier = pipeline(
|
| 59 |
-
|
|
|
|
|
|
|
|
|
|
| 60 |
)
|
| 61 |
-
|
|
|
|
| 62 |
print(predictions)
|
| 63 |
```
|
| 64 |
|
|
|
|
|
|
|
| 65 |
## Training data
|
| 66 |
|
| 67 |
Some of the entities (for instance WORK_OF_ART, LAW, MONEY) that have been annotated in the [Turku OntoNotes Entities Corpus](https://github.com/TurkuNLP/turku-one)
|
|
|
|
| 52 |
The easiest way to use the model is by utilizing the Transformers pipeline for token classification:
|
| 53 |
|
| 54 |
```python
|
| 55 |
+
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
|
| 56 |
|
| 57 |
model_checkpoint = "Kansallisarkisto/finbert-ner"
|
| 58 |
+
tokenizer = AutoTokenizer.from_pretrained(model_checkpoint)
|
| 59 |
+
model = AutoModelForTokenClassification.from_pretrained(model_checkpoint)
|
| 60 |
+
|
| 61 |
token_classifier = pipeline(
|
| 62 |
+
"token-classification",
|
| 63 |
+
model=model,
|
| 64 |
+
tokenizer=tokenizer,
|
| 65 |
+
aggregation_strategy="simple",
|
| 66 |
)
|
| 67 |
+
|
| 68 |
+
predictions = token_classifier("Helsingistä tuli Suomen suuriruhtinaskunnan pääkaupunki vuonna 1812.")
|
| 69 |
print(predictions)
|
| 70 |
```
|
| 71 |
|
| 72 |
+
Running the code requires installing Transformers and PyTorch libraries.
|
| 73 |
+
|
| 74 |
## Training data
|
| 75 |
|
| 76 |
Some of the entities (for instance WORK_OF_ART, LAW, MONEY) that have been annotated in the [Turku OntoNotes Entities Corpus](https://github.com/TurkuNLP/turku-one)
|