gliner2-PII-basque-v2

Version 2 of tknika/gliner2-PII-basque, a Basque-adapted fine-tune of fastino/gliner2-privacy-filter-PII-multi (205M-parameter multilingual PII detection model, GLiNER2 architecture, mDeBERTa-v3-base encoder).

Compared to v1, this version is trained with a v2 synthetic replay dataset (tknika/pii-synthetic-basque-v2) with realistic value distributions (name frequencies from public statistics, coherent street/town/postcode triples from OpenStreetMap) and three entity types specific to educational contexts: user name (LMS/forum handles, including @-mentions), personal url (personal blogs and profiles) and student id.

Training

Two data sources were combined (experience replay, to avoid catastrophic forgetting of the base model's PII capabilities):

  • Basque NER: the nerc_id split of orai-nlp/basqueGLUE (2,842 sentences), mapped to the base model's types: person name, location, organization, miscellaneous.
  • Synthetic PII: the train split (5,000 sentences) of tknika/pii-synthetic-basque-v2 — Basque, Spanish and French sentences covering the whole Basque Country (Araba, Bizkaia, Gipuzkoa, Nafarroa, Iparralde), 17 PII types, checksum-valid values (mod-23 national IDs, mod-97 IBANs, Luhn-valid cards, real phone formats).

Training used LoRA (r=16, alpha=32) on the boundary/task heads, merged into the base model after training (~13 MB adapter). This repository contains the merged standalone model, directly loadable with AutoExtractor.from_pretrained(); no additional files or merging steps are needed.

Results

PII detection

On the eval split of pii-synthetic-basque-v2 (1,000 sentences, 17 types, disjoint from training by construction; exact (type, mention) matching):

Model Precision Recall Micro F1 False positives on negatives
Base (zero-shot) 0.680 0.794 0.733 397
v1 (gliner2-PII-basque) 0.774 0.882 0.824 61
This model (v2) 0.960 0.990 0.975 32

Per-type F1 of this model (types sorted by frequency; base / v1 shown for the educational types):

Type F1 Type F1
person name 0.957 personal url 0.954 (base 0.149, v1 0.670)
email 0.997 organization 0.948
date 1.000 location 0.812
user name 0.994 (base 0.846, v1 0.615) student id 1.000 (base 0.871, v1 0.937)
phone number 1.000 bank account number 1.000
address 0.995 national id 0.978
city 0.922 age 1.000
date of birth 1.000 credit card number 1.000
zip code 1.000

Basque NER

BasqueGLUE validation sets (500 sentences per split), showing that the PII fine-tuning does not cause catastrophic forgetting of the Basque NER capabilities:

Metric Base (zero-shot) This model (v2)
nerc_id/val person name F1 0.664 0.808
nerc_id/val micro F1 0.464 0.706
nerc_od/val (Wikipedia) person name F1 0.721 0.762
nerc_od/val micro F1 0.547 0.697

Usage

from gliner2 import AutoExtractor

model = AutoExtractor.from_pretrained("tknika/gliner2-PII-basque-v2")

text = ("Kaixo, @mikel_etxeberria naiz ikaslea, ikasle zk 123456B dut. "
        "Nire bloga https://mikeletxeberria.wordpress.com da eta "
        "helbidea Barakaldoko Nagusia kalea 42 da, 48901.")
result = model.extract_entities(
    text,
    ["person name", "email", "phone number", "national id",
     "bank account number", "credit card number", "date of birth",
     "date", "address", "city", "location", "zip code", "age",
     "organization", "user name", "personal url", "student id"],
)

Limitations

  • The PII evaluation set is synthetic; real-world performance (messier text, ambiguous contexts) will be lower.
  • Some residual city/location confusion remains (city F1 0.922, location F1 0.812): town names are occasionally labelled as location and provinces as city.
  • The synthetic replay data covers 17 of the 42 PII types of the base model; other types may degrade after fine-tuning.
  • Research-quality evaluation; not production-tested.
  • The license of the BasqueGLUE/EIEC training corpus should be verified before redistribution of derivatives beyond this model.

Citation

If you use this model, please cite:

@misc{tknika2026gliner2piibasquev2,
  title  = {gliner2-PII-basque-v2: A Basque-Adapted Multilingual PII Detection Model with Educational Entity Types},
  author = {{TKNIKA} and Ezpeleta Mendikute, Xabier},
  year   = {2026},
  url    = {https://huggingface.co/tknika/gliner2-PII-basque-v2}
}

This model builds on the following work, which you may also want to cite:

@misc{fastino2026gliner2pii,
  title   = {GLiNER2-PII: Multilingual PII Extraction via Synthetic Fine-Tuning},
  author  = {{Fastino AI Team}},
  year    = {2026},
  url     = {https://huggingface.co/fastino/gliner2-pii-v1}
}

@misc{zaratiana2026gliner2piimultilingualmodelpersonally,
  title  = {GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction},
  author = {Zaratiana, Urchade and Lewis, Ash and Hurn-Maloney, George},
  year   = {2026},
  eprint = {2605.09973},
  archivePrefix = {arXiv},
  url   = {https://arxiv.org/abs/2605.09973}
}

@InProceedings{urbizu2022basqueglue,
  author    = {Urbizu, Gorka and San Vicente, Iñaki and Saralegi, Xabier and Agerri, Rodrigo and Soroa, Aitor},
  title     = {BasqueGLUE: A Natural Language Understanding Benchmark for Basque},
  booktitle = {Proceedings of the Language Resources and Evaluation Conference},
  year      = {2022},
  address   = {Marseille, France},
  publisher = {European Language Resources Association},
  pages     = {1603--1612},
  url       = {https://aclanthology.org/2022.lrec-1.172}
}

@misc{he2021debertav3,
  title  = {DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing},
  author = {He, Pengcheng and Gao, Jianfeng and Chen, Weizhu},
  year   = {2021},
  eprint = {2111.09543},
  archivePrefix = {arXiv},
  url   = {https://arxiv.org/abs/2111.09543}
}
Downloads last month
34
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tknika/gliner2-PII-basque-v2

Finetuned
(5)
this model

Datasets used to train tknika/gliner2-PII-basque-v2

Papers for tknika/gliner2-PII-basque-v2

Evaluation results