PMxbench leaderboard
PMxbench · a pharmacometrics benchmark from the AIML-SIG
Leaderboard
Every scored attempt at a PMxbench scenario. Further right, and darker, is better. A score of 1 means the answers matched the hidden answer key.
- 28 scored runs
- 10 configurations
- Scenario 00
How a run gets a score
data.csv and the analysis plan.submission.yamlEach run is scored from 0 to 1 in four areas, and the overall score rolls them into one number. A configuration is one combination of tool, harness (the agent program) and model. baseline is one plain agent call with no scaffolding, the floor to compare against.
Overall score, every run
One dot per run, the bar is the mean. Rows are sorted by mean. Dots that land on top of each other are nudged apart so you can count them.
- baselinecodex / gpt-5.6-terra0.943
- moduspi / z-ai/glm-5.20.929
- baselinepi / z-ai/glm-5.20.859
- baselineclaude / sonnet0.845
- modusclaude / opus0.843
- baselineclaude / opus0.832
- modusclaude / sonnet0.826
- moduscodex / gpt-5.6-terra0.726
- modusclaude / haiku0.207
- baselineclaude / haiku0.161
Where each configuration gets it right
Mean score by pharmacometric area. Read across a row to see which parts of the analysis a configuration handles well, and which it misses.
- StructuralDid the analysis pick the right number of compartments?
- EstimationHow close are the estimated PK parameters and residual error to the true values?
- Covariate analysisDid it keep the covariates that really matter, on the right parameters, and drop the rest?
- Data qualityDid it notice the records in the data that should not be trusted?
- OverallOne headline number for the run, across all four areas.
| Configuration | Structural | Estimation | Covariate analysis | Data quality | Overall |
|---|---|---|---|---|---|
| baselinecodex / gpt-5.6-terra | 1.000 | 0.929 | 1.000 | 0.857 | 0.943 |
| moduspi / z-ai/glm-5.2 | 1.000 | 0.888 | 0.868 | 1.000 | 0.929 |
| baselinepi / z-ai/glm-5.2 | 1.000 | 0.773 | 0.749 | 1.000 | 0.859 |
| baselineclaude / sonnet | 1.000 | 0.718 | 0.789 | 1.000 | 0.845 |
| modusclaude / opus | 1.000 | 0.723 | 0.770 | 1.000 | 0.843 |
| baselineclaude / opus | 1.000 | 0.798 | 0.563 | 1.000 | 0.832 |
| modusclaude / sonnet | 1.000 | 0.715 | 0.698 | 1.000 | 0.826 |
| moduscodex / gpt-5.6-terra | 1.000 | 0.420 | 0.791 | 1.000 | 0.726 |
| modusclaude / haiku | 0.333 | 0.030 | 0.083 | 0.556 | 0.207 |
| baselineclaude / haiku | 0.000 | 0.083 | 0.106 | 0.536 | 0.161 |
Scores are recomputed from each stored submission.yaml whenever this page is built, so every run is held to the current answer key.