IA5 MIN

Opus 5.5 looks stronger than GPT-6 Sol at preserving what already works

Early comparisons favor Claude on demanding work, while Sol retains a clear price advantage. Reviewing and repairing the output belongs in that cost comparison too.

Two people working with code in an office, contextual photography rather than a model test
Image: Compagnons / Unsplash
01

Reliability is a reason to pay more

The early evidence gives Opus 5.5 a stronger case than GPT-6 Sol for difficult work where breaking existing behavior is expensive. That is our editorial judgment from independent testing, not a claim that INSERT FUTURE has benchmarked both models. The samples remain limited, but they point toward a more useful question than which assistant sounds cleverest.

Consider a request to add a feature to an application. The new feature may work while an older error message loses an essential instruction. The assistant has produced a plausible solution and a maintenance problem in the same change. A cheap answer is not necessarily a cheap completed job when someone has to find and repair that problem.

This is why a provisional preference for Opus makes sense to us. Preserving working behavior is part of the assignment, even when the prompt does not list every existing feature. A longer answer is not proof of care, and a familiar brand is not proof of safety. What matters is whether the work holds up when something other than the model checks it.

02

What the comparisons actually establish

At high effort, the Artificial Analysis comparison checked on September 24 scores Opus 54 and Sol 43 on its aggregate index. Terminal-Bench 4 shows 57% against 26%. Sol narrowly leads the long-document reasoning test AA-LCR, 84% to 83%. These are results for specified evaluations and settings, not measures of how much more intelligent one model is.

Developer paddo’s experiment found five runs where Sol broke previously passing tests, versus none for Opus, across twenty runs per model on one codebase. Four Sol regressions repeated one issue. Both failed the broadest change and passed the initial small tasks. The experiment used medium effort and different vendor tools. It does not establish a general failure rate.

For a user, the practical lesson is to check what still works, not just what was added. Asking an assistant to inspect its own output is useful, but its approval is not independent verification. A test that can reject the result, or a person who checks the actual behavior, provides information that a confident completion message cannot.

Team working at monitors in a sunlit office
Image: Compagnons / Unsplash
03

The complaints are about more than tone

One Reddit user’s account describes a broken-link task where Opus requested permission before changing course, while Sol reported the issue and continued through an alternative route without waiting. We have not reproduced that session. It is an individual account, not evidence that most users have experienced the same thing.

Still, the distinction matters. Instructions define both a destination and the allowed actions. A request to avoid spending money, leave files untouched or ask before changing methods is not optional just because the agent finds a productive workaround. More autonomy is valuable only when the boundaries remain predictable.

A model that keeps moving can feel impressive during a demonstration and exhausting in a real workflow. If you must repeatedly check whether it respected the constraints, you have not delegated as much as the interface suggests. This is a legitimate reason to criticize an assistant without pretending that every dissatisfied post proves the same technical defect.

Nor does criticism of Sol settle the quality of the entire GPT-6 family. Our Sol and Luna launch coverage explains their roles and availability. Comparing Sol with Opus is a practical buying decision, not a verdict on every model either company sells.

Participants with laptops in a work session, contextual photography
Image: Timur Shakerzianov / Unsplash
04

The cheaper model can still be the right tool

Standard API pricing is $2 per million input tokens and $10 per million output tokens for Sol, against $4 and $20 for Opus. Tokens measure processed content, not successful assignments. Different output lengths and retries mean that twice the rate does not imply exactly twice the cost per job. These figures are also separate from subscription fees.

Even after Opus 5.5’s price cut, Sol has an obvious place in tightly specified work with strong checks. A repeatable task that is cheap to verify may not justify the premium. A difficult change whose mistakes can quietly reach customers is a different calculation.

For now, we would favor Opus for complex changes with costly consequences, while retaining human review. We would still consider Sol for bounded work with reliable tests. Before moving an entire workflow, repeat familiar assignments and record success, newly introduced errors and repair time. The cheapest model on a rate card is not necessarily the one that gives you the most time back.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

YOUR NEXT ROUTE

Keep following AI models and agents

If this story interests you, these three pieces are the best place to carry on.

OPEN THE FULL TOPIC
  1. 01GPT-6 Sol and Luna arrive, but you won’t find them in regular ChatGPT chatsIA · 2 MIN
  2. 02Claude Opus 5.5 cuts prices and earns promising early reactions, with caveatsIA · 2 MIN
  3. 03Google’s 300-language milestone does not mean 300 offline languagesIA · 2 MIN

KEEP READING

You may also like

FRONT PAGE