PMxbench leaderboard

PMxbench · a pharmacometrics benchmark from the AIML-SIG

Leaderboard

Every scored attempt at a PMxbench scenario. Further right, and darker, is better. A score of 1 means the answers matched the hidden answer key.

  • 28 scored runs
  • 10 configurations
  • Scenario 00
Take part

How a run gets a score

Data and plan
data.csv and the analysis plan.
Analysis
By hand, by an AI agent, or both.
submission.yaml
Compared with the hidden answer key.

Each run is scored from 0 to 1 in four areas, and the overall score rolls them into one number. A configuration is one combination of tool, harness (the agent program) and model. baseline is one plain agent call with no scaffolding, the floor to compare against.

Overall score, every run

One dot per run, the bar is the mean. Rows are sorted by mean. Dots that land on top of each other are nudged apart so you can count them.

  1. baselinecodex / gpt-5.6-terra1 run
    One run: 0.943Mean of 1 run: 0.943
    0.943
  2. moduspi / z-ai/glm-5.23 runs
    One run: 0.955One run: 0.986One run: 0.846Mean of 3 runs: 0.929, spread ±0.074
    0.929
  3. baselinepi / z-ai/glm-5.25 runs
    One run: 0.845One run: 0.946One run: 0.824One run: 0.948One run: 0.731Mean of 5 runs: 0.859, spread ±0.091
    0.859
  4. baselineclaude / sonnet3 runs
    One run: 0.849One run: 0.838One run: 0.848Mean of 3 runs: 0.845, spread ±0.006
    0.845
  5. modusclaude / opus3 runs
    One run: 0.839One run: 0.846One run: 0.846Mean of 3 runs: 0.843, spread ±0.004
    0.843
  6. baselineclaude / opus2 runs
    One run: 0.716One run: 0.948Mean of 2 runs: 0.832, spread ±0.164
    0.832
  7. modusclaude / sonnet4 runs
    One run: 0.835One run: 0.776One run: 0.852One run: 0.840Mean of 4 runs: 0.826, spread ±0.034
    0.826
  8. moduscodex / gpt-5.6-terra1 run
    One run: 0.726Mean of 1 run: 0.726
    0.726
  9. modusclaude / haiku3 runs
    One run: 0.462One run: 0.024One run: 0.134Mean of 3 runs: 0.207, spread ±0.227
    0.207
  10. baselineclaude / haiku3 runs
    One run: 0.262One run: 0.202One run: 0.021Mean of 3 runs: 0.161, spread ±0.125
    0.161

Where each configuration gets it right

Mean score by pharmacometric area. Read across a row to see which parts of the analysis a configuration handles well, and which it misses.

  • StructuralDid the analysis pick the right number of compartments?
  • EstimationHow close are the estimated PK parameters and residual error to the true values?
  • Covariate analysisDid it keep the covariates that really matter, on the right parameters, and drop the rest?
  • Data qualityDid it notice the records in the data that should not be trusted?
  • OverallOne headline number for the run, across all four areas.
Configuration Structural Estimation Covariate analysis Data quality Overall
baselinecodex / gpt-5.6-terra1 run 1.000 0.929 1.000 0.857 0.943
moduspi / z-ai/glm-5.23 runs 1.000 0.888 0.868 1.000 0.929
baselinepi / z-ai/glm-5.25 runs 1.000 0.773 0.749 1.000 0.859
baselineclaude / sonnet3 runs 1.000 0.718 0.789 1.000 0.845
modusclaude / opus3 runs 1.000 0.723 0.770 1.000 0.843
baselineclaude / opus2 runs 1.000 0.798 0.563 1.000 0.832
modusclaude / sonnet4 runs 1.000 0.715 0.698 1.000 0.826
moduscodex / gpt-5.6-terra1 run 1.000 0.420 0.791 1.000 0.726
modusclaude / haiku3 runs 0.333 0.030 0.083 0.556 0.207
baselineclaude / haiku3 runs 0.000 0.083 0.106 0.536 0.161

Scores are recomputed from each stored submission.yaml whenever this page is built, so every run is held to the current answer key.