The brief Sol got fitted in one paragraph
On 9 July, OpenAI researcher Aidan McLaughlin posted something on X that two years ago would have led a news bulletin: «i cannot tell you how routine it is for me to have 5.6 e2e do an entire rl run». Routine, he says. A full reinforcement learning run, set up and supervised by the model itself, mentioned in the tone of somebody noting they have put a wash on.
The Decoder unpacked what sits behind that post the following day, and it still reads to me like the biggest thing that has happened this month: GPT-5.6 Sol, the large model of the house, handled the post-training of Luna, the small one. The brief arrived through Codex and OpenAI describes it as «fairly underspecified» — find the right training configurations, pick suitable GPUs, launch the training script, and verify everything was running correctly. From there, on its own. When we covered the launch of the three models, this part had not been told yet.
OpenAI's Kathy Shi put it plainly: «previously this is something that a team of senior researchers may have worked on at OpenAI, and now it really feels like the automated researcher is pretty close».
On the internal index the company uses to measure exactly that — recursive self-improvement, or how far one model can get you towards building the next — Sol scores 16.2 points above GPT-5.5. And that jump is already part of how the place works day to day.
The student costs a fifth of what the teacher costs
This is where it gets interesting for whoever pays the bill. Sol charges $5 per million input tokens and $30 per million output. Terra, the middle tier, sits at $2.50 and $15. And Luna, the model that came out of that brief, runs at $1 in and $6 out: exactly one fifth of the big one on both sides.
The gap in results does not follow that ratio, nowhere near. On TerminalBench 2.1, the terminal-agent test that has become the industry's thermometer, the table Wikipedia compiles gives Sol 88.8% and Luna 82.5%. Six points and change. The figure that made me raise an eyebrow sits one row further down: GPT-5.5, the flagship of five minutes ago, lands on 83.4%. Which means Luna, at the price it charges, is nine tenths of a point off the best model OpenAI knew how to build in spring.
The small print is still where it always was, mind. Those rates apply to contexts under 270,000 tokens; cache writes are billed at 1.25x the uncached input rate, though cache reads come in 90% cheaper and batch processing knocks off half. None of it is outrageous, but it is worth a look before you build an agent that spends its day re-reading the same context.
The same slide, handed to both
Percentages get tiring, so a concrete job is worth more. In the launch material OpenAI gave both models the same task: here is a corporate slide, match its style and build me another one with this data. No coding involved, an office job of the kind you knock out on a Thursday afternoon.
GPT-5.5 came back with the chart below. It is fine: the bars are right, the axes read cleanly and the data sits where it should. It could also have come from anywhere. The brand typeface, the corporate colour, the white card the original was mounted on and the value labels above each bar all fell off along the way.

GPT-5.6 copies the legal note at the bottom too
And this is what GPT-5.6 returned from the same prompt. The brand template is reproduced whole, white card, colour band, value labels above every bar and even the tiny disclaimer in the footer. Stacked one above the other, the first looks like a classroom exercise and the second like a slide you could send without touching it.
OpenAI sums up its own pitch in a sentence: «more intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work». It reads like a brochure, because it is one, but underneath there is a number Altman has repeated more than once: a 54% gain in token efficiency on agentic coding work. Fewer tokens for the same job means, simply, a smaller bill at the end of the month.

Altman has started using the present tense for the singularity
Inside OpenAI this has stopped looking like an experiment. In its own announcement the company says internal agentic token usage has grown roughly 22-fold in six months, that compute allocated to internal coding inference is up 100x, and that average daily output tokens per active researcher have run at more than twice the highest level they had ever seen. Weekly experiments per researcher have doubled since January.
With that in the background, the line Sam Altman dropped a few days ago on the Relentless podcast — «we are now, like, in the singularity» — reads differently than it would have a year ago. In his June 2025 essay he called it «a larval version of recursive self-improvement». Thirteen months on, his big model configures and supervises the training of the small one, and the result sells at a dollar per million input tokens.
The brake comes from the same place. On 14 July, Altman warned in Axios that Sol's growth «is insane», that they are going to «move mountains» to keep scaling, and that even so «it is possible there are some hiccups soon». An industry that already trains itself and still cannot keep up with a Tuesday's worth of requests.
For now, the one who answers for whatever goes wrong is still a person.
The conversation starts here
Sign in with a supporter account to comment. Sign in



Nobody has commented yet. Want to go first?