Public experiment results

LEPISZCZE
Leaderboard.

Compare Polish language models across the complete benchmark — including the extractive question answering datasets developed through Expansio.

evaluations
dataset views
models

Explore results

Rank models.
Inspect every metric.

Rankings are recalculated in your browser. Choose a task, dataset and ranking metric; click any metric header to sort the table directly.

Loading verified results…

Reproducible by design

Results with an audit trail.

The public dataset contains only model identity, dataset metadata and aggregate metrics. Source commit and schema are validated during every Pages deployment. AITAX and JuDDGES datasets are tracked separately in the dataset pipeline until they pass the publication, validation and reproducibility gates required for ranking.