Which subscription does the work?
Atris Bench gives paid AI tools the same practical work, then publishes the result and the transcript behind it.
Leaderboard
Ranked by reliable pass, then first-attempt pass rate, then speed.
| Rank | Model and subscription | Reliable pass3 | Pass at one | Median time | Tasks won | Transcripts |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1Claude Max · $200/mo | 100% | 100% | 19s | 5 | 24 |
| 2 | Muse Spark 1.3opencode (contributor free) · $0/mo | 100% | 100% | 24s | 2 | 24 |
| 3 | GPT-5.6 Sol (Codex)ChatGPT Pro · $200/mo | 100% | 100% | 34s | 0 | 24 |
| 4 | Grok 4.6SuperGrok · $30/mo | 100% | 100% | 35s | 0 | 24 |
| 5 | Gemini 3.8 Flash (High)Antigravity · $0/mo | 75% | 87.5% | 52s | 0 | 24 |
Task results
Open any scored cell to read the best transcript for that model and task.
| Model | Find the payout mismatches | Triage a small inbox | Fix two invoicing bugs | Find shared calendar slots | Remove dead JavaScript exports | Reply to a dental customer | Write the owner's morning brief | Answer contract questions |
|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Muse Spark 1.3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| GPT-5.6 Sol (Codex) | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Grok 4.6 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 | 3/3 |
| Gemini 3.8 Flash (High) | 3/3 | 3/3 | 3/3 | 3/3 | 2/3 | 2/3 | 3/3 | 3/3 |
How it works
The tasks are ordinary business work, including reconciling payouts, triaging inboxes, fixing small bugs, and reading contracts.
Each model runs through the agent command line tool included with the subscription a buyer would pay for.
Every model receives the same fresh folder and the same prompt, with tools available and no hints.
Code checks decide the result where possible. When judgment is needed, a model from a rival company judges the work.
Reliable pass measures whether every attempt passed. Pass at one shows how often the first attempt worked.
Every result links to a transcript. New task packs arrive weekly, and old packs remain available for trend checks.
What this does not measure
Atris Bench does not measure raw API quality, long-horizon work, or cost per token. These products sell flat-rate subscriptions, so the comparison shows completion time instead of token cost.