🔍 Read the full analysis: How I Organize AI Work With Opus, Sol, And Jev on ThorstenMeyerAI.com
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
On 29 September 2026, Thorsten Meyer published a cost-driven AI workflow: Claude Opus 5.5 at high or xhigh effort as the main builder, the newly released GPT-6.1 Sol as a low-cost reviewer, and Jev for high-volume routing decisions. With six frontier models within roughly 20 index points but 100x apart in cost per task, model selection shifts from capability to cost-efficiency.
AI analyst Thorsten Meyer published a working model stack on 29 September 2026 that routes daily AI work across Claude Opus 5.5, the newly released GPT-6.1 Sol, and Jev, a decision-only model, arguing that with six frontier models clustered within about 20 index points of each other but roughly 100x apart in cost per task, the practical question has shifted from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?”
The core of the workflow is a two-model pairing. Opus 5.5, released 22 September, is the main builder, run at high effort for features, APIs, and multi-file work ($1.82 per task, 54 points on the Artificial Analysis Intelligence Index v4.3.x) and at xhigh for architecture, migrations, and trust boundaries ($3.46 per task, 56 points). GPT-6.1 Sol, released the day of publication, serves as the second pair of eyes: at high or xhigh effort it costs $0.32 to $0.39 per task for scores of 50 to 51, making routine review passes affordable on every meaningful change.
Meyer’s comparison table, built on Artificial Analysis Intelligence Index v4.3.x data, shows the spread: Opus 5.5 tops the field at 58 points ($5.98 per task at max), while GPT-6 Luna sits at 37 points but costs $0.07 per task — 1,429 tasks per $100. Between them, Claude Sonnet 5.5 (56 points, $7.60 at max), Claude Fable 5.1 (53, $7.63), GPT-6 Astra (53, $3.26), and Sol fill specific niches across the frontier model lineup. Meyer notes that Sonnet 5.5 at max effort costs more per task than Opus at max while scoring 2 points lower, and that Opus now outscores its more expensive sibling Fable by 5 points at lower cost.
The effort setting emerges as the biggest cost lever. On Opus 5.5, moving from xhigh to max adds 2 index points for 73% more cost per task; from medium to max, cost rises 4.46x for 7 points. Sol has real catches, per Meyer: high and xhigh settings take 57 to 69 seconds to first token, ruling out interactive use, and Opus still leads it by 5 points at xhigh.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost-Per-Task Now Drives Model Choice
The report captures a structural change in how AI work gets done. When capability gaps between frontier models shrink to single index points while prices differ by orders of magnitude, the marginal value of “the best model” collapses for most tasks. Meyer’s example makes the point: Sol xhigh sits 1-2 points under Astra and Fable at $0.39 instead of $3.26 or $7.63 per task — meaning a review pass costs little enough to run routinely rather than selectively.
The cross-family review pattern is the second takeaway. A different model family reviewing Opus’s output is, in Meyer’s framing, a better check than Opus reviewing itself — but he adds the caveat that a second model reading the same flawed spec is not an independent review. The workflow also assigns Jev, a decision model that cannot write a sentence, to high-volume yes/no and routing judgements, extending the cost logic below the text-generation tier.
A Month of Rapid Frontier Releases
: “The stack rests on an unusually dense release window. In September 2026 alone: Fable 5.1 (1 September), Astra (3 September), GPT-6 Luna and Opus 5.5 (22 September), Sonnet 5.5 (28 September), and GPT-6.1 Sol (29 September). Sol launched at the same $2/$10 per 1M tokens as its week-old predecessor; even its medium setting matches the earlier GPT-6 Sol’s score of 48 at one-fifth of that model’s $1.06 per task, per Artificial Analysis data cited by Meyer.
Meyer grounds his choices in four stated rules: effort is not capability; more effort cannot fill in missing requirements; passing tests are not approval to ship; and failing cases get handed to the builder with evidence, never just “try harder.”
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve. Six models now sit within about 20 index points of each other, while their cost per task differs by roughly 100×.”
— Thorsten Meyer
Limits of the Benchmark and the Math
Meyer is explicit that his scores come from one source — the Artificial Analysis Intelligence Index v4.3.x — and that one index point is inside the noise. Artificial Analysis has not yet published low or max settings for Sol, so the full cost curve for that model is incomplete.
The cost-of-ownership math is also qualified by the author himself. His observation that halving model price saves only 12.5% of real cost, and that a single extra minute of human review can erase the saving, is described as illustrative, not measured. Whether the Opus-plus-Sol split generalizes beyond Meyer’s own development workload is untested; he recommends shadow-testing before any switch.
Watching Sol’s Missing Settings
The immediate open item is Artificial Analysis publishing low and max effort settings for GPT-6.1 Sol, which would complete its cost curve and show whether cheaper settings preserve usable quality. Meyer indicates he will keep Opus at high/xhigh as his default and reserve max settings for rare cases, while running Sol review passes on every meaningful change and using Astra or Fable as tie-breakers only when Sol and Opus disagree. Whether vendors respond to the price compression with their own cuts, and whether the September release cadence continues into October, remains to be seen.
Key Questions
What is the core workflow described?
Opus 5.5 at high or xhigh effort builds software and handles knowledge work; GPT-6.1 Sol at $0.32-$0.39 per task performs detail dives and independent review; Jev, a decision model, handles high-volume yes/no and routing calls.
Why not just use the highest-scoring model for everything?
According to Meyer, the top settings rarely justify their cost. Opus 5.5 max adds only 2 points over xhigh for 73% more cost per task, and Sonnet 5.5 at max costs more than Opus at max while scoring lower.
What are GPT-6.1 Sol’s main drawbacks?
High and xhigh settings take 57 to 69 seconds to produce a first token, making it unsuitable for interactive use, and Opus 5.5 still leads it by 5 index points at xhigh (56 vs. 51).
Are the benchmark scores a reliable guide for my own work?
Meyer says no — the Artificial Analysis index maps general capability, not any specific workload, and one index point is within measurement noise. He recommends shadow-testing models before switching.
Does a cheaper model actually reduce total costs?
Only partially. Meyer’s illustrative (not measured) example holds that halving model price saves about 12.5% of real cost, an amount a single extra minute of human review can erase.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
