🔍 Read the full analysis: Opus, Sol, And Jev: A Practical AI Workflow For September on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
With six frontier AI models now within about 20 index points of each other while per-task costs differ by roughly 100x, Thorsten Meyer proposes a workflow built on Claude Opus 5.5 as main builder, newly released GPT-6.1 Sol as reviewer, and the Jev decision model for high-volume routing. All benchmarks are from the Artificial Analysis Intelligence Index v4.3.x.
GPT-6.1 Sol launched on 29 September 2026, and independent AI commentator Thorsten Meyer used the release to publish a practical model workflow that pairs Claude Opus 5.5 as the main building model with Sol as a low-cost reviewer and the Jev decision model for high-volume routing. The recommendation rests on a simple observation: six frontier models now sit within roughly 20 points of each other on the Artificial Analysis Intelligence Index v4.3.x, while their cost per task differs by about 100x — turning model choice from a quality question into a cost-per-task question.
At the top of the index, Opus 5.5 scores 58 at its max effort setting and costs $5.98 per task (17 tasks per $100). GPT-6.1 Sol at xhigh scores 51 for $0.39 per task — 256 tasks per $100 — and GPT-6 Luna scores 37 for $0.07 per task (1,429 per $100). Between those poles sit Claude Sonnet 5.5 (56 at max, $7.60), Claude Fable 5.1 (53, $7.63) and GPT-6 Astra (53, $3.26). Meyer reports three findings from the data: Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task; Sonnet 5.5 at max effort costs more than Opus at max for 2 fewer points; and Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
Meyer argues the effort setting, not model choice, is the dominant cost lever. On Opus 5.5, moving from xhigh to max adds 2 index points for 73% more cost per task; from medium to max, cost rises 4.46x for 7 points. His recommended operating points: Opus at high (54 points, $1.82 per task) for features and multi-file work, and xhigh (56, $3.46) for architecture, migrations and trust boundaries. Max is, in his assessment, rarely worth it. Sonnet 5.5 at max produces about 193k output tokens per task — the most Artificial Analysis has measured, according to Meyer — with cost jumping from $2.74 to $7.60 for 4 points.
GPT-6.1 Sol launched at the same $2/$10 per 1M tokens as its week-old predecessor. Artificial Analysis already lists three effort levels: medium (48, $0.21), high (50, $0.32) and xhigh (51, $0.39). Meyer notes Sol is unusually concise — 25M output tokens on the index at high, against a median of 82M for comparable models — but that high and xhigh take 57 to 69 seconds to the first token, ruling it out for interactive use. The index has not yet published low or max settings for Sol, and Meyer cautions that one index point falls inside measurement noise.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost Per Task Now Beats Leaderboards
The workflow reflects a broader shift in how practitioners choose frontier models. When scores cluster within roughly 20 points but prices span two orders of magnitude, the deciding question becomes which model clears a given quality bar at the lowest cost per task — not which model tops a leaderboard. For teams running agents, review pipelines or bulk classification at volume, the price gaps Meyer documents translate directly into budget: at $0.39 per task, a Sol review pass on every meaningful change is affordable as routine practice, an approach that would be uneconomical at Opus or Fable pricing.
The structure of Meyer’s stack also encodes a review principle: a model from a different family checking Opus’s output is a stronger check than Opus reviewing itself, he argues, and cheap review tokens make that second seat viable at scale. He adds four caveats to keep the arrangement honest: effort is not capability, more effort cannot fill in missing requirements, a different model is not independent if both read the same flawed spec, and passing tests is not approval to ship.
The September Release Cluster Behind It
The analysis covers a dense four-week release window. Claude Fable 5.1 shipped 1 September, GPT-6 Astra on 3 September, GPT-6 Luna and Opus 5.5 on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September — the same day Meyer published. All scores cited come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability rather than a verdict on any specific workload. Per-1M-token pricing he lists: Opus 5.5 at $4/$20 (cache reads $0.20), Fable and Astra at $10/$50, Sol at $2/$10, and Luna at $0.10/$0.50.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve. Six models now sit within about 20 index points of each other, while their cost per task differs by roughly 100x.”
— Thorsten Meyer, ThorstenMeyerAI.com
What the Index Does Not Yet Show
Several elements remain open. Artificial Analysis has not published low or max effort settings for GPT-6.1 Sol, and Meyer notes that single index points fall within measurement noise, so gaps of 1 to 2 points between Sol, Astra and Fable should not be treated as decisive. The index is a general-capability measure, not a workload-specific one; Meyer repeatedly advises shadow-testing before switching any production model. His cost-versus-human-review example is, by his own label, illustrative rather than measured. The Jev decision model — which Meyer says handles high-volume yes/no and routing judgements and cannot write a sentence — is named in the workflow but not benchmarked in the published data, so its cost and accuracy figures are not provided.
Watching Sol’s Full Curve and Jev’s Benchmarks
The near-term items to watch are the low and max effort results for GPT-6.1 Sol once Artificial Analysis publishes them, which could change its value calculus at either end of the curve. Sol’s time-to-first-token of 57 to 69 seconds at high and xhigh may improve in later iterations, which would determine whether it can move from a background review role into interactive use. Meyer’s workflow also implies continued reliance on Jev for routing and classification at volume, but its published benchmarks are absent so far. Readers following the stack should treat the 29 September configuration as a snapshot of a fast-moving field, and re-verify index versions and prices before acting on any of it.
Key Questions
What is the core recommendation of this workflow?
Use Claude Opus 5.5 at high or xhigh effort as the main model for building software, use GPT-6.1 Sol for detail work and independent review passes, reserve Astra and Fable for second opinions, use Sonnet 5.5 for scoped subtasks, and route high-volume yes/no judgements to the Jev decision model.
Why is GPT-6.1 Sol described as a reviewer rather than a builder?
According to Meyer, Sol scores 51 at xhigh — about 5 points below Opus 5.5 — but costs roughly $0.32 to $0.39 per task, one-eighth of Astra’s and one-twentieth of Fable’s cost. That makes it cheap enough to run on every meaningful change. Its 57-69 second time to first token at higher effort settings also rules out interactive use.
What is the ‘effort setting’ cost lever?
Effort levels (low, medium, high, xhigh, max) change both score and cost. On Opus 5.5, going from xhigh to max adds 2 index points but 73% more cost per task; medium to max costs 4.46x more for 7 points. Meyer calls this a bigger cost factor than most model swaps.
Where do the benchmark numbers come from?
All scores come from the Artificial Analysis Intelligence Index v4.3.x, unless otherwise noted. Meyer cautions that the index measures general capability, not any specific workload, and recommends shadow-testing before switching models.
What is Jev’s role in the stack?
Jev is a decision model that cannot generate sentences, used for high-volume yes/no calls and routing judgements. Meyer assigns it classification and extraction work alongside GPT-6 Luna, though no benchmark figures for Jev appear in the published data.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
