HukukBERT: Domain-Specific Language Model for Turkish Law

Authors

  • Mehmet Utku Öztürk Kalitte Inc., Türkiye
  • Tansu Türkoğlu Aibrite Inc., Türkiye
  • Buse Buz-Yalug University of Eastern Finland, Finland

DOI:

https://doi.org/10.47852/bonviewJCLLT620210346

Keywords:

Turkish legal NLP, domain-adaptive pre-training, legal language model, WordPiece tokenization, Legal Cloze Test

Abstract

Natural language processing (NLP) advances have powered a generation of LegalTech systems, but Turkish law remains under-served by domain-specific data and models. English has legal encoders such as LEGAL-BERT; no comparable high-volume Turkish counterpart exists. We introduce HukukBERT, a Turkish legal language model trained on a 19 GB cleaned corpus using a hybrid domain-adaptive pre-training (DAPT) recipe that mixes Whole‐Word Masking, Token Span Masking, Word Span Masking, and targeted Keyword Masking. We compared our 48K WordPiece tokenizer and DAPT pipeline against general-purpose and existing domain-specific Turkish models. On the Legal Cloze Test—a masked legal term prediction benchmark over Turkish court decisions—HukukBERT reaches 84.40% Top-1 accuracy and beats every baseline we tested. The Legal Cloze Test is synthetically constructed, so its passages are absent from the pre-training corpus by construction, eliminating train–test contamination. On the downstream task of structural segmentation of official Turkish court decisions, it reaches a 92.8% document pass rate. We release HukukBERT to support Turkish legal NLP work in named entity recognition, judgment prediction, and document classification.

 

Received: 9 May 2026 | Revised: 25 June 2026 | Accepted: 13 July 2026

 

Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.


Data Availability Statement

The Hukuki Cloze Testi (Legal Cloze Test) benchmark is openly available on the Hugging Face Hub at https://huggingface.co/datasets/turkhukuk/hukukbert-cloze-benchmark. The associated evaluation scripts for the benchmark are available at https://github.com/TurkHukuk/hukukbert. The pre-trained HukukBERT model weights and the training code are not publicly released by the authors at this time, though they may be shared privately for research collaboration purposes.


Author Contribution Statement

Mehmet Utku Öztürk: Conceptualization, Software, Data curation, Writing – original draft, Writing – review & editing, Visualization. Tansu Türkoğlu: Conceptualization, Software, Data curation, Resources. Buse Buz-Yalug: Methodology, Validation, Writing – review & editing, Supervision.

Downloads

Published

2026-08-04

Issue

Section

Research Articles

How to Cite

Öztürk, M. U., Türkoğlu, T., & Buz-Yalug, B. (2026). HukukBERT: Domain-Specific Language Model for Turkish Law. Journal of Computational Law and Legal Technology, 1-15. https://doi.org/10.47852/bonviewJCLLT620210346