AtrisPricing

Which subscription does the work?

Atris Bench gives paid AI tools the same practical work, then publishes the result and the transcript behind it.

Leaderboard

Ranked by reliable pass, then first-attempt pass rate, then speed.

Subscription benchmark leaderboard
RankModel and subscriptionReliable pass3Pass at oneMedian timeTasks wonTranscripts
1Claude Fable 5.1Claude Max · $200/mo100%100%19s524
2Muse Spark 1.3opencode (contributor free) · $0/mo100%100%24s224
3GPT-5.6 Sol (Codex)ChatGPT Pro · $200/mo100%100%34s024
4Grok 4.6SuperGrok · $30/mo100%100%35s024
5Gemini 3.8 Flash (High)Antigravity · $0/mo75%87.5%52s024

Task results

Open any scored cell to read the best transcript for that model and task.

Passes and attempts for every model and task
ModelFind the payout mismatchesTriage a small inboxFix two invoicing bugsFind shared calendar slotsRemove dead JavaScript exportsReply to a dental customerWrite the owner's morning briefAnswer contract questions
Claude Fable 5.13/33/33/33/33/33/33/33/3
Muse Spark 1.33/33/33/33/33/33/33/33/3
GPT-5.6 Sol (Codex)3/33/33/33/33/33/33/33/3
Grok 4.63/33/33/33/33/33/33/33/3
Gemini 3.8 Flash (High)3/33/33/33/32/32/33/33/3

How it works

The tasks are ordinary business work, including reconciling payouts, triaging inboxes, fixing small bugs, and reading contracts.

Each model runs through the agent command line tool included with the subscription a buyer would pay for.

Every model receives the same fresh folder and the same prompt, with tools available and no hints.

Code checks decide the result where possible. When judgment is needed, a model from a rival company judges the work.

Reliable pass measures whether every attempt passed. Pass at one shows how often the first attempt worked.

Every result links to a transcript. New task packs arrive weekly, and old packs remain available for trend checks.

What this does not measure

Atris Bench does not measure raw API quality, long-horizon work, or cost per token. These products sell flat-rate subscriptions, so the comparison shows completion time instead of token cost.

Generated Sep 4, 2026, 12:03 PM Coordinated Universal Time · Pack business-v1