Public experiment results
LEPISZCZE
Leaderboard.
Compare Polish language models across the complete benchmark — including the extractive question answering datasets developed through Expansio.
Explore results
Rank models.
Inspect every metric.
Rankings are recalculated in your browser. Choose a task, dataset and ranking metric; click any metric header to sort the table directly.
—
—
—
No models match this search.
Values show the mean over repeated evaluations; denotes standard deviation where available. Higher values rank first for all published metrics.
This is a minimal, immutable snapshot of the LEPISZCZE leaderboard. The QA datasets were incorporated after the 2022 paper; full experiment configurations remain in the source repository. Inspect source ↗
Reproducible by design
Results with an audit trail.
The public dataset contains only model identity, dataset metadata and aggregate metrics. Source commit and schema are validated during every Pages deployment. AITAX and JuDDGES datasets are tracked separately in the dataset pipeline until they pass the publication, validation and reproducibility gates required for ranking.