Instructions to use tknika/gliner2-PII-basque-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use tknika/gliner2-PII-basque-v2 with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque-v2") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
gliner2-PII-basque-v2
Version 2 of tknika/gliner2-PII-basque, a Basque-adapted fine-tune of fastino/gliner2-privacy-filter-PII-multi (205M-parameter multilingual PII detection model, GLiNER2 architecture, mDeBERTa-v3-base encoder).
Compared to v1, this version is trained with a v2 synthetic replay dataset (tknika/pii-synthetic-basque-v2) with realistic value distributions (name frequencies from public statistics, coherent street/town/postcode triples from OpenStreetMap) and three entity types specific to educational contexts: user name (LMS/forum handles, including @-mentions), personal url (personal blogs and profiles) and student id.
Training
Two data sources were combined (experience replay, to avoid catastrophic forgetting of the base model's PII capabilities):
- Basque NER: the
nerc_idsplit of orai-nlp/basqueGLUE (2,842 sentences), mapped to the base model's types:person name,location,organization,miscellaneous. - Synthetic PII: the train split (5,000 sentences) of tknika/pii-synthetic-basque-v2 — Basque, Spanish and French sentences covering the whole Basque Country (Araba, Bizkaia, Gipuzkoa, Nafarroa, Iparralde), 17 PII types, checksum-valid values (mod-23 national IDs, mod-97 IBANs, Luhn-valid cards, real phone formats).
Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training (~13 MB adapter). This repository contains the merged standalone model, directly loadable with AutoExtractor.from_pretrained(); no additional files or merging steps are needed.
Results
PII detection
On the eval split of pii-synthetic-basque-v2 (1,000 sentences, 17 types, disjoint from training by construction; exact (type, mention) matching):
| Model | Precision | Recall | Micro F1 | False positives on negatives |
|---|---|---|---|---|
| Base (zero-shot) | 0.680 | 0.794 | 0.733 | 397 |
| v1 (gliner2-PII-basque) | 0.774 | 0.882 | 0.824 | 61 |
| This model (v2) | 0.960 | 0.990 | 0.975 | 32 |
Per-type F1 of this model (types sorted by frequency; base / v1 shown for the educational types):
| Type | F1 | Type | F1 | |
|---|---|---|---|---|
| person name | 0.957 | personal url | 0.954 (base 0.149, v1 0.670) | |
| 0.997 | organization | 0.948 | ||
| date | 1.000 | location | 0.812 | |
| user name | 0.994 (base 0.846, v1 0.615) | student id | 1.000 (base 0.871, v1 0.937) | |
| phone number | 1.000 | bank account number | 1.000 | |
| address | 0.995 | national id | 0.978 | |
| city | 0.922 | age | 1.000 | |
| date of birth | 1.000 | credit card number | 1.000 | |
| zip code | 1.000 |
Basque NER
BasqueGLUE validation sets (500 sentences per split), showing that the PII fine-tuning does not cause catastrophic forgetting of the Basque NER capabilities:
| Metric | Base (zero-shot) | This model (v2) |
|---|---|---|
| nerc_id/val person name F1 | 0.664 | 0.808 |
| nerc_id/val micro F1 | 0.464 | 0.706 |
| nerc_od/val (Wikipedia) person name F1 | 0.721 | 0.762 |
| nerc_od/val micro F1 | 0.547 | 0.697 |
Usage
from gliner2 import AutoExtractor
model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque-v2")
text = ("Kaixo, @mikel_etxeberria naiz ikaslea, ikasle zk 123456B dut. "
"Nire bloga https://mikeletxeberria.wordpress.com da eta "
"helbidea Barakaldoko Nagusia kalea 42 da, 48901.")
result = model.extract_entities(
text,
["person name", "email", "phone number", "national id",
"bank account number", "credit card number", "date of birth",
"date", "address", "city", "location", "zip code", "age",
"organization", "user name", "personal url", "student id"],
)
Limitations
- The PII evaluation set is synthetic; real-world performance (messier text, ambiguous contexts) will be lower.
- Some residual
city/locationconfusion remains (city F1 0.922, location F1 0.812): town names are occasionally labelled aslocationand provinces ascity. - The synthetic replay data covers 17 of the 42 PII types of the base model; other types may degrade after fine-tuning.
- Research-quality evaluation; not production-tested.
- The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.
Citation
If you use this model, please cite:
@misc{tknika2026gliner2piibasquev2,
title = {gliner2-PII-basque-v2: A Basque-Adapted Multilingual PII Detection Model with Educational Entity Types},
author = {{TKNIKA} and Ezpeleta Mendikute, Xabier},
year = {2026},
url = {https://huggingface.co/tknika/gliner2-PII-basque-v2}
}
This model builds on the following work, which you may also want to cite:
@misc{fastino2026gliner2pii,
title = {GLiNER2-PII: Multilingual PII Extraction via Synthetic Fine-Tuning},
author = {{Fastino AI Team}},
year = {2026},
url = {https://huggingface.co/fastino/gliner2-pii-v1}
}
@misc{zaratiana2026gliner2piimultilingualmodelpersonally,
title = {GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction},
author = {Zaratiana, Urchade and Lewis, Ash and Hurn-Maloney, George},
year = {2026},
eprint = {2605.09973},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.09973}
}
@InProceedings{urbizu2022basqueglue,
author = {Urbizu, Gorka and San Vicente, Iñaki and Saralegi, Xabier and Agerri, Rodrigo and Soroa, Aitor},
title = {BasqueGLUE: A Natural Language Understanding Benchmark for Basque},
booktitle = {Proceedings of the Language Resources and Evaluation Conference},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {1603--1612},
url = {https://aclanthology.org/2022.lrec-1.172}
}
@misc{he2021debertav3,
title = {DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing},
author = {He, Pengcheng and Gao, Jianfeng and Chen, Weizhu},
year = {2021},
eprint = {2111.09543},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2111.09543}
}
- Downloads last month
- 34
Model tree for tknika/gliner2-PII-basque-v2
Base model
fastino/gliner2-privacy-filter-PII-multiDatasets used to train tknika/gliner2-PII-basque-v2
tknika/pii-synthetic-basque-v2
Papers for tknika/gliner2-PII-basque-v2
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
Evaluation results
- Micro F1 on pii-synthetic-basque-v2 (eval split)self-reported0.975
- Micro Precision on pii-synthetic-basque-v2 (eval split)self-reported0.960
- Micro Recall on pii-synthetic-basque-v2 (eval split)self-reported0.990
- Micro F1 on basqueGLUE (nerc_id/validation)validation set self-reported0.706
- Person name F1 on basqueGLUE (nerc_id/validation)validation set self-reported0.808