Legal NER: Evaluating the Impact of LLM-Generated Annotations on NER Performance for Administrative Decisions

Authors

  • Harry Nan Department of Public Law and Governance, Tilburg University, The Netherlands
  • Samaneh Khoshrou Department of Intelligent Systems, Tilburg University, The Netherlands
  • Johan Wolswinkel Department of Public Law and Governance, Tilburg University, The Netherlands

DOI:

https://doi.org/10.47852/bonviewJCLLT62029445

Keywords:

named entity recognition, information extraction, relationship extraction, large language models, annotation

Abstract

Named entity recognition (NER) is a core information extraction (IE) task dependent on high-quality annotated data that are expensive and time-intensive to produce. Large language models (LLMs) offer a promising alternative through LLM-generated pseudo-annotations, yet their reliability in domain-specific legal settings remains insufficiently studied. This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language. To this end, LLM-based annotations of predefined legal entities are created using a schema-driven few-shot prompt, which are evaluated against a human-annotated dataset. Next, two different NER architectures are trained—a token-level NER model and a span-based NER–RE model (joint NER and relationship extraction (RE))—under three training settings: (1) human annotations only, (2) LLM-generated annotations only, and (3) models trained on LLM-generated annotations further fine-tuned on human annotations. The results indicate that LLMs can accurately generate annotations for legal entities that are explicitly defined in legislation but generate less reliable annotations for other legal entities that require deeper contextual understanding beyond what is explicitly stated in the text or prompt. Furthermore, fine-tuning a NER model trained on these LLM-generated annotations with human annotations turns out to slightly outperform models trained on human-annotated data only. Our findings highlight the potential of hybrid supervision strategies to scale low-resource legal NER tasks while maintaining human-level accuracy.

 

Received: 23 February 2026 | Revised: 29 June 2026 | Accepted: 23 July 2026

 

Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.


Data Availability Statement

The data that support the findings of this study are openly available on GitHub.


Author Contribution Statement

Harry Nan: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – original draft, Writing – review & editing, Visualization, Project administration. Samaneh Khoshrou: Conceptualization, Validation, Resources, Writing – review & editing, Supervision, Project administration. Johan Wolswinkel: Conceptualization, Validation, Resources, Data curation, Writing – review & editing, Supervision, Project administration, Funding acquisition.

Downloads

Published

2026-08-10

Issue

Section

Special Issue: Translating Natural Legal Language into Formal Representation

How to Cite

Nan, H., Khoshrou, S., & Wolswinkel, J. (2026). Legal NER: Evaluating the Impact of LLM-Generated Annotations on NER Performance for Administrative Decisions. Journal of Computational Law and Legal Technology, 1-13. https://doi.org/10.47852/bonviewJCLLT62029445

Funding data