IA6 MIN

OpenAI's priciest model trained its cheapest one. Luna costs a fifth of Sol and lands six points behind

OpenAI says Sol picked the training configurations, chose the GPUs, launched the run and checked it was working. Internal agent token usage is up 22x in six months. Altman has started using the present tense for the singularity.

OpenAI's official illustration of the Sun, Earth and Moon to scale on a starfield, labelled with the three GPT-5.6 model names
Image: OpenAI
01

The brief Sol got fitted in one paragraph

On 9 July, OpenAI researcher Aidan McLaughlin posted something on X that two years ago would have led a news bulletin: «i cannot tell you how routine it is for me to have 5.6 e2e do an entire rl run». Routine, he says. A full reinforcement learning run, set up and supervised by the model itself, mentioned in the tone of somebody noting they have put a wash on.

The Decoder unpacked what sits behind that post the following day, and it still reads to me like the biggest thing that has happened this month: GPT-5.6 Sol, the large model of the house, handled the post-training of Luna, the small one. The brief arrived through Codex and OpenAI describes it as «fairly underspecified» — find the right training configurations, pick suitable GPUs, launch the training script, and verify everything was running correctly. From there, on its own. When we covered the launch of the three models, this part had not been told yet.

OpenAI's Kathy Shi put it plainly: «previously this is something that a team of senior researchers may have worked on at OpenAI, and now it really feels like the automated researcher is pretty close».

On the internal index the company uses to measure exactly that — recursive self-improvement, or how far one model can get you towards building the next — Sol scores 16.2 points above GPT-5.5. And that jump is already part of how the place works day to day.

02

The student costs a fifth of what the teacher costs

This is where it gets interesting for whoever pays the bill. Sol charges $5 per million input tokens and $30 per million output. Terra, the middle tier, sits at $2.50 and $15. And Luna, the model that came out of that brief, runs at $1 in and $6 out: exactly one fifth of the big one on both sides.

The gap in results does not follow that ratio, nowhere near. On TerminalBench 2.1, the terminal-agent test that has become the industry's thermometer, the table Wikipedia compiles gives Sol 88.8% and Luna 82.5%. Six points and change. The figure that made me raise an eyebrow sits one row further down: GPT-5.5, the flagship of five minutes ago, lands on 83.4%. Which means Luna, at the price it charges, is nine tenths of a point off the best model OpenAI knew how to build in spring.

The small print is still where it always was, mind. Those rates apply to contexts under 270,000 tokens; cache writes are billed at 1.25x the uncached input rate, though cache reads come in 90% cheaper and batch processing knocks off half. None of it is outrageous, but it is worth a look before you build an agent that spends its day re-reading the same context.

03

The same slide, handed to both

Percentages get tiring, so a concrete job is worth more. In the launch material OpenAI gave both models the same task: here is a corporate slide, match its style and build me another one with this data. No coding involved, an office job of the kind you knock out on a Thursday afternoon.

GPT-5.5 came back with the chart below. It is fine: the bars are right, the axes read cleanly and the data sits where it should. It could also have come from anywhere. The brand typeface, the corporate colour, the white card the original was mounted on and the value labels above each bar all fell off along the way.

Slide generated by GPT-5.5: a plain light-blue bar chart with no brand template and no value labels
Image: OpenAI
05

Altman has started using the present tense for the singularity

Inside OpenAI this has stopped looking like an experiment. In its own announcement the company says internal agentic token usage has grown roughly 22-fold in six months, that compute allocated to internal coding inference is up 100x, and that average daily output tokens per active researcher have run at more than twice the highest level they had ever seen. Weekly experiments per researcher have doubled since January.

With that in the background, the line Sam Altman dropped a few days ago on the Relentless podcast — «we are now, like, in the singularity» — reads differently than it would have a year ago. In his June 2025 essay he called it «a larval version of recursive self-improvement». Thirteen months on, his big model configures and supervises the training of the small one, and the result sells at a dollar per million input tokens.

The brake comes from the same place. On 14 July, Altman warned in Axios that Sol's growth «is insane», that they are going to «move mountains» to keep scaling, and that even so «it is possible there are some hiccups soon». An industry that already trains itself and still cannot keep up with a Tuesday's worth of requests.

For now, the one who answers for whatever goes wrong is still a person.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE