AI models and agents
New models, tool-using agents and open projects: what they do well, where they fall short and why they matter.
Benchmarks with context, concrete tests and explanations that do not require memorising every previous release.READING ROUTE
THE DOSSIER
Start here and follow the story.
01OPENAlibaba built an AI that 'imagines' what a computer will do next—and uses it to train better agents
Qwen-AgentWorld simulates terminals, websites, search, Android and MCP tools so other agents can practise without touching live systems. Its training results are promising, although the benchmark placing it above GPT-5.4 needs important context.
02OPENMoonshot's Kimi K3 rattled Silicon Valley—but its open-weight promise is still due July 27
China's 2.8-trillion-parameter model can handle a million-token context and has impressed coding evaluators. It is already reviving DeepSeek-era anxiety in the US, even though the weights, licence and technical report have not arrived.
03OPENGPT-5.6 is here, and it wants to do the whole job—not just answer questions
OpenAI is launching a model built to research, use tools and complete long assignments. It could save hours, but tracing a mistake across all those steps will be harder.
04OPENCurrent AI wants a public ChatGPT—and its first demo already works
Alpha Chat already works without an account, shows what it uses to answer and deletes conversations after 30 days. It is only version 0.1, but the coalition behind it has more than $400 million pledged and wants to turn open AI into a public service.