The leaderboard UI now brings filter leaders, animated benchmark expansion, Intelligence Index fit scores, and paged release notes into one clearer workspace.
A leaderboard for sub-150M parameter language models, evaluated using LM-eval harness or a custom benchmark script available here ArithMark-3.0.
To support our work and help us keep this leaderboard up to date, please consider giving the space a like and follow!
Short notes on new models and leaderboard shifts.
The leaderboard UI now brings filter leaders, animated benchmark expansion, Intelligence Index fit scores, and paged release notes into one clearer workspace.
ArithMark-3 is now the primary math benchmark for ranking, weighting subject to change. ArithMark-2 remains available in the expanded columns.
UCR's 2.74M SLM uses custom tokenization and digit features to reach 69.4% ArithMark accuracy, setting the leaderboard's #1 score
Built on Axiomic Labs' T-X4 refresh-gated XSA stack, GPT-S2-5M pushes into the sub-10M bracket and lands at the top, dethroning SLM-10M.
Veyra AI's 49M Apricot checkpoint posts a 37.63 average, setting the strongest score in the sub-50M bracket.
Zero-shot evaluation. Higher is better for all columns. Click any header to sort.
| Compare | # ▼ | Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Challenge | PIQA | ArithMark-3 | Avg | Fit Std Devs | ArithMark-2 |
|---|
Models you add from the leaderboard line up side by side here. None selected
Top scores for the active size and benchmark filters.
Score vs parameter count (log scale). Shaded zone = above regression line.
Average standard deviations above or below the Intelligence Index-vs-size fit line.
| # | Organization | Models | Fit Std Devs | Mean Int Index | Best Model vs Fit |
|---|
Eligibility, evaluation, and submission requirements for the Open SLM Leaderboard.
The leaderboard is for language models with fewer than 150 million parameters.
The submitted model must be pretrained from scratch by you or your organization.
Model weights must be openly available and point to a genuine training checkpoint, not a merged model.
Provide zero-shot results for the listed benchmarks using LM-eval harness or the linked custom ArithMark script.
Submitted results are independently checked by the leaderboard team before they are merged.
Open a PR or discussion with the benchmark results in the leaderboard Space.
Open a PR or discussion on this Space with your model's results for the given benchmarks. They will be independently verified by our team and then your PR will be merged. Your model must be pretrained by you from scratch, open weights and be a training checkpoint (not merged) to qualify Open a PR →
Each benchmark is first adjusted for its random-chance floor, so chance performance maps to 0 and perfect performance maps to 100.
Combined ARC is the mean of ARC-Easy and ARC-Challenge before normalization. When a component is unavailable, its weight is removed from both the numerator and denominator; at least two components are required. Scores below chance may be negative. ArithMark-2 is displayed separately and is not included in the index.