NeurIPS 2022 · Datasets & Benchmarks

Polish NLP,
measured with context.

LEPISZCZE is a comprehensive benchmark designed to evaluate language models across diverse Polish language understanding tasks — with reproducible experiments and a blueprint for other underrepresented languages.

Published inNeurIPS 35
LanguagePolish
TrackingDVC + W&B
LicenseMIT

Why LEPISZCZE

One benchmark.
A wider field of view.

Progress in language models should be visible across more than one narrow task. LEPISZCZE combines established Polish benchmarks with new datasets to test how models generalize across domains, structures and prediction problems.

02

Reproducible

Versioned pipelines, fixed experiment configurations and tracked runs make the evaluation process inspectable rather than opaque.

03

Extensible

Models, datasets and tasks are represented as configuration-driven pipeline stages, lowering the cost of extending the benchmark.

Benchmark coverage

Tasks grounded in real Polish text.

The benchmark spans legal language, reviews, news, Wikipedia, social media, general-domain corpora and extractive question answering. All current datasets are listed below, including the Expansio QA datasets maintained as part of LEPISZCZE.

LEPISZCZE · Question answering

Extractive QA, inside LEPISZCZE.

Three Polish SQuAD 2.0-style datasets are part of the evolving LEPISZCZE benchmark, adding answerable, paraphrased and unanswerable questions. Developed through Expansio, they joined LEPISZCZE after the original 2022 publication.

Compare QA results

Dataset pipeline · Legal AI

Research data,
with visible maturity.

LEPISZCZE tracks datasets developed through AITAX and JuDDGES before they enter the official benchmark. Public corpora, previews and benchmark candidates remain clearly separated from comparable leaderboard results.

tracked datasets
benchmark candidates
2contributing projects
Benchmark candidate

Public data with a defined task and a visible path to reproducible evaluation.

Public preview

Useful for research, but metadata, annotations or the evaluation protocol may still change.

Public dataset

A reusable corpus or derived asset that is not presented as a leaderboard test set.

Closest to evaluation

Featured datasets

Each card states why the dataset matters and which gates still prevent official ranking.

Loading dataset registry…

Complete registry

AITAX and JuDDGES collections

Promotion gate

A dataset joins the official leaderboard only after its version and test split are frozen, licensing is complete, leakage risks are reviewed, metrics and baselines are published, and the evaluation can be reproduced from a versioned protocol.

Project resources

Go from paper
to pipeline.

Read the methodology, inspect versioned experiment definitions, explore public datasets and review the published evaluation results.

Reproducibility note

Pipeline definitions are public. Historical DVC outputs use a non-public remote due to artifact size; contact the authors if you need access to those generated artifacts.

Use the benchmark

Cite LEPISZCZE.

If this benchmark supports your research, please cite the NeurIPS 2022 paper.

@inproceedings{augustyniak2022lepiszcze,
  author    = {Augustyniak, Lukasz and Tagowski, Kamil and Sawczyn, Albert
               and Janiak, Denis and Bartusiak, Roman and Szymczak, Adrian
               and Janz, Arkadiusz and Szymański, Piotr and Wątroba, Marcin
               and Morzy, Mikołaj and Kajdanowicz, Tomasz and Piasecki, Maciej},
  title     = {This is the way: designing and compiling LEPISZCZE,
               a comprehensive NLP benchmark for Polish},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {35},
  pages     = {21805--21818},
  publisher = {Curran Associates, Inc.},
  year      = {2022}
}