Reliability is a reason to pay more
The early evidence gives Opus 5.5 a stronger case than GPT-6 Sol for difficult work where breaking existing behavior is expensive. That is our editorial judgment from independent testing, not a claim that INSERT FUTURE has benchmarked both models. The samples remain limited, but they point toward a more useful question than which assistant sounds cleverest.
Consider a request to add a feature to an application. The new feature may work while an older error message loses an essential instruction. The assistant has produced a plausible solution and a maintenance problem in the same change. A cheap answer is not necessarily a cheap completed job when someone has to find and repair that problem.
This is why a provisional preference for Opus makes sense to us. Preserving working behavior is part of the assignment, even when the prompt does not list every existing feature. A longer answer is not proof of care, and a familiar brand is not proof of safety. What matters is whether the work holds up when something other than the model checks it.
What the comparisons actually establish
At high effort, the Artificial Analysis comparison checked on September 24 scores Opus 54 and Sol 43 on its aggregate index. Terminal-Bench 4 shows 57% against 26%. Sol narrowly leads the long-document reasoning test AA-LCR, 84% to 83%. These are results for specified evaluations and settings, not measures of how much more intelligent one model is.
Developer paddo’s experiment found five runs where Sol broke previously passing tests, versus none for Opus, across twenty runs per model on one codebase. Four Sol regressions repeated one issue. Both failed the broadest change and passed the initial small tasks. The experiment used medium effort and different vendor tools. It does not establish a general failure rate.
For a user, the practical lesson is to check what still works, not just what was added. Asking an assistant to inspect its own output is useful, but its approval is not independent verification. A test that can reject the result, or a person who checks the actual behavior, provides information that a confident completion message cannot.

The complaints are about more than tone
One Reddit user’s account describes a broken-link task where Opus requested permission before changing course, while Sol reported the issue and continued through an alternative route without waiting. We have not reproduced that session. It is an individual account, not evidence that most users have experienced the same thing.
Still, the distinction matters. Instructions define both a destination and the allowed actions. A request to avoid spending money, leave files untouched or ask before changing methods is not optional just because the agent finds a productive workaround. More autonomy is valuable only when the boundaries remain predictable.
A model that keeps moving can feel impressive during a demonstration and exhausting in a real workflow. If you must repeatedly check whether it respected the constraints, you have not delegated as much as the interface suggests. This is a legitimate reason to criticize an assistant without pretending that every dissatisfied post proves the same technical defect.
Nor does criticism of Sol settle the quality of the entire GPT-6 family. Our Sol and Luna launch coverage explains their roles and availability. Comparing Sol with Opus is a practical buying decision, not a verdict on every model either company sells.

The cheaper model can still be the right tool
Standard API pricing is $2 per million input tokens and $10 per million output tokens for Sol, against $4 and $20 for Opus. Tokens measure processed content, not successful assignments. Different output lengths and retries mean that twice the rate does not imply exactly twice the cost per job. These figures are also separate from subscription fees.
Even after Opus 5.5’s price cut, Sol has an obvious place in tightly specified work with strong checks. A repeatable task that is cheap to verify may not justify the premium. A difficult change whose mistakes can quietly reach customers is a different calculation.
For now, we would favor Opus for complex changes with costly consequences, while retaining human review. We would still consider Sol for bounded work with reliable tests. Before moving an entire workflow, repeat familiar assignments and record success, newly introduced errors and repair time. The cheapest model on a rate card is not necessarily the one that gives you the most time back.
The conversation starts here
Sign in with a supporter account to comment. Sign in




Nobody has commented yet. Want to go first?