AI models and agents
New models, tool-using agents and open projects: what they do well, where they fall short and why they matter.
Benchmarks with context, concrete tests and explanations that do not require memorizing every previous release.READING ROUTE
THE DOSSIER
Start here and follow the story.
01OPENThe new Copilot wants to do the work for you. Here is what launches now and what costs extra
Microsoft is bringing chat, app creation and persistent agents into Home, Code and Autopilot. Access starts in Frontier, while longer agent tasks will use consumption-based billing.
02OPENGGUF or safetensors: which file should you download to run AI locally?
Choose the format your inference app supports. A file extension alone does not tell you whether a model is smarter, smaller or compatible with your computer.
03OPENGemini 3.8 Live improves voice chat, but screen sharing deserves a closer look
Google adds a reasoning variant to its voice-model family. Before showing it an app, understand the separate controls for microphone, camera and screen sharing.
04OPENSetting AI temperature to zero does not make its answers true
Temperature changes how a model selects tokens. It can reduce variation, but it does not check sources, and some models work best with their default settings.
05OPENContext caching can lower AI costs. It does not add memory or make answers free
Reusing a long document can avoid repeated input processing through an API. Savings depend on cache hits, storage charges and the output you still generate.
06OPENClaude can use your private files in Slack. The team may still see its answer
Claude Tag now lets people use personal connectors inside Slack channels. The access remains theirs, but any answer posted to the channel becomes visible to its members.
07OPENMeta wants Muse to see what you see and help book your next night out
Connect brings a clearer picture of Meta’s personal agent: visual context from glasses, more shopping connections and live-event bookings. Several of the most useful features are still on the way.
08OPENOpus 5.5 looks stronger than GPT-6 Sol at preserving what already works
Early comparisons favor Claude on demanding work, while Sol retains a clear price advantage. Reviewing and repairing the output belongs in that cost comparison too.
09OPENGPT-6 Sol and Luna arrive, but you won’t find them in regular ChatGPT chats
OpenAI adds two GPT-6 models for coding and work. Access depends on the product, and early users are watching usage limits as closely as the answers.
10OPENClaude Sonnet 5.5 gets close to Opus. The real surprise is what it costs to run
Anthropic's new model makes a striking leap in coding without raising its token price. Independent testing finds a catch: pushing it to its strongest setting can cost more per task than Sonnet 5.
11OPENClaude Opus 5.5 cuts prices and earns promising early reactions, with caveats
Anthropic’s claimed saving per task is different from its per-token price cut. Early users welcome the speed, while others still report mistakes.
12OPENGoogle’s 300-language milestone does not mean 300 offline languages
The headline figure spans Google’s technologies. TranslateGemma offers a route to local translation, not a 300-language offline pack for your phone.
12 of 79 stories showing
