Broad by design
Classification, sequence labeling, entailment, sentiment, paraphrase, named entities, part-of-speech tagging, punctuation restoration and extractive question answering in one evaluation framework.
NeurIPS 2022 · Datasets & Benchmarks
LEPISZCZE is a comprehensive benchmark designed to evaluate language models across diverse Polish language understanding tasks — with reproducible experiments and a blueprint for other underrepresented languages.
Why LEPISZCZE
Progress in language models should be visible across more than one narrow task. LEPISZCZE combines established Polish benchmarks with new datasets to test how models generalize across domains, structures and prediction problems.
Classification, sequence labeling, entailment, sentiment, paraphrase, named entities, part-of-speech tagging, punctuation restoration and extractive question answering in one evaluation framework.
Versioned pipelines, fixed experiment configurations and tracked runs make the evaluation process inspectable rather than opaque.
Models, datasets and tasks are represented as configuration-driven pipeline stages, lowering the cost of extending the benchmark.
Benchmark coverage
The benchmark spans legal language, reviews, news, Wikipedia, social media, general-domain corpora and extractive question answering. All current datasets are listed below, including the Expansio QA datasets maintained as part of LEPISZCZE.
Image captions
Wikipedia
Online reviews
News
Legal texts
Online reviews
LEPISZCZE · Question answering
Three Polish SQuAD 2.0-style datasets are part of the evolving LEPISZCZE benchmark, adding answerable, paraphrased and unanswerable questions. Developed through Expansio, they joined LEPISZCZE after the original 2022 publication.
Compare QA resultsDataset pipeline · Legal AI
LEPISZCZE tracks datasets developed through AITAX and JuDDGES before they enter the official benchmark. Public corpora, previews and benchmark candidates remain clearly separated from comparable leaderboard results.
Public data with a defined task and a visible path to reproducible evaluation.
Useful for research, but metadata, annotations or the evaluation protocol may still change.
A reusable corpus or derived asset that is not presented as a leaderboard test set.
Closest to evaluation
Each card states why the dataset matters and which gates still prevent official ranking.
Complete registry
A dataset joins the official leaderboard only after its version and test split are frozen, licensing is complete, leakage risks are reviewed, metrics and baselines are published, and the evaluation can be reproduced from a versioned protocol.
Project resources
Read the methodology, inspect versioned experiment definitions, explore public datasets and review the published evaluation results.
Pipeline definitions are public. Historical DVC outputs use a non-public remote due to artifact size; contact the authors if you need access to those generated artifacts.
Use the benchmark
If this benchmark supports your research, please cite the NeurIPS 2022 paper.
@inproceedings{augustyniak2022lepiszcze,
author = {Augustyniak, Lukasz and Tagowski, Kamil and Sawczyn, Albert
and Janiak, Denis and Bartusiak, Roman and Szymczak, Adrian
and Janz, Arkadiusz and Szymański, Piotr and Wątroba, Marcin
and Morzy, Mikołaj and Kajdanowicz, Tomasz and Piasecki, Maciej},
title = {This is the way: designing and compiling LEPISZCZE,
a comprehensive NLP benchmark for Polish},
booktitle = {Advances in Neural Information Processing Systems},
volume = {35},
pages = {21805--21818},
publisher = {Curran Associates, Inc.},
year = {2022}
}