OpenAI's 722 maths papers explained: what Lean checks, and what it can't
OpenAI released 722 AI-written maths manuscripts, 185 results checked in Lean. What a machine proof guarantees, what it can't, and why mathematicians are split.
Mistral's 1-trillion-parameter Large 4 'Le Chonk' is in preview, weights due late October. How its sparse design works, what it scores, and what is unproven.
Mistral has shown the model that a cat meme predicted. On 6 October the Paris company opened a public preview of Mistral Large 4, a roughly one-trillion-parameter model it calls "Le Chonk", trained from scratch in its own European data centres, with downloadable weights promised by the end of October (27 October, according to VentureBeat). The benchmarks are a mixed picture, but the argument Mistral is making is sharper than the scores: that in a world where a US provider can withdraw or restrict a model, a copy you hold yourself is the only one you can plan around.
Think of a large language model as a very big reference library with a librarian. When you ask a question, the librarian does not read every book. They walk to the two or three shelves that matter and read those. Mistral Large 4 is built this way. It holds about 1.05 trillion "parameters", the numbers that encode what it has learned, which is the whole library. But for each word it produces it consults only about 50 billion of them, the few relevant shelves. That design is called a "mixture of experts", and it is why a model this large can answer at a speed Artificial Analysis measured at about 116 words-worth of tokens per second. The catch is that the library still has to exist. You cannot run the model on a laptop; you need a rack of server chips holding roughly a terabyte of memory.
Mistral's announcement, published on its site, makes three claims. First, that the model was built entirely in Europe, on 3,800 NVIDIA Grace Blackwell chips in Mistral's own data centres, trained from scratch rather than copied from someone else's model. Second, that it is the strongest model you will be able to download that was built outside China. Third, that it is unusually good at computer security work, like finding a flaw in software and writing a fix.
Or jump to your depth:
Why it matters: The cheapest tier of frontier-lab models is where most high-volume agent and classification work runs, and a tenfold price cut on small prompts resets what that work costs. It also sharpens the price war with OpenAI's Luna line and Mistral's new Large 4.
Why it matters: This is the first time GPT-6 reaches ChatGPT's free tier. The interface change matters as much as the model: chat answers that build their own mini-apps move ChatGPT closer to a software platform than a text box.
Why it matters: Microsoft is betting that agents will live on the PC rather than only in the cloud, and the sandbox is the piece that makes that defensible for IT departments. Local models with no metered token cost also change the economics of coding assistants.
Microsoft Windows Developer Blog · NVIDIA · X · Tom's Hardware
Why it matters: The AI buildout is shifting from corporate bonds to bespoke private-credit and lease structures, which keep debt off balance sheets but spread the risk across more lenders. Morgan Stanley estimates the sector needs $1.5 trillion in outside financing by 2028.
Reuters via Yahoo Finance · investingLive, citing The Wall Street Journal
Why it matters: Memory is now the bottleneck and the profit centre of the AI hardware stack. Record margins for Samsung mean higher component costs for everyone else who builds on DRAM and HBM.
Why it matters: This is the first large independent test of OpenAI's teen product since it launched in August, and it lands as regulators in the US and EU tighten rules on minors and chatbots. The parental-alert finding is the one OpenAI will have to answer most directly.
Why it matters: At more than 6 GW across two days, this is one of the largest power commitments any AI company has made, and it leans on existing reactors rather than new builds. Expect rivals to compete for the same limited pool of uprate capacity.
Why it matters: Europe's data-centre boom is running into permitting law, and this is a rare case of a regulator stopping a hyperscaler mid-build. A delay at Google's largest European investment is a warning for other operators clearing land ahead of approvals.
Why it matters: A Musk-run product defaulting to a rival's frontier model is an admission that Grok is not the best engine for agent work. It also deepens an unusual relationship: Anthropic already rents compute in SpaceX's Colossus data centre.
Why it matters: Google is aiming AI generation at user-made games, Roblox's home turf, with distribution through the browser and Google Play Games profiles. It is also a consumer showcase for stacking several Google models in one product.
Why it matters: Open-source agents are becoming venture-scale businesses by selling the governance layer rather than the model. Microsoft naming Hermes among the tools adopting its new Windows agent sandbox shows how fast that layer is being standardised.
TechCrunch · Runtime Wire · X
Why it matters: It arrives weeks after Amazon blocked Meta's Muse agent from shopping on its store, and it overlaps with Visa's Trusted Agent Protocol, which Stripe and Shopify have also joined. Whoever sets the sign-in standard controls which agents retailers trust.
Why it matters: It is the largest model a European lab has promised to open, and Mistral is pitching it at governments wanting sovereign AI. Until the weights and licence land, it is a closed API model competing on price.
Mistral AI · Decrypt · Stork
OpenAI released 722 AI-written maths manuscripts, 185 results checked in Lean. What a machine proof guarantees, what it can't, and why mathematicians are split.
OpenAI will watermark ChatGPT and Codex text in the EU with textGrain. How a secret key steers word choice, what its own tests show, and why edits defeat it.
Trump has launched a Super Intelligence Force led by intelligence chief Jay Clayton. What it is, what it cannot do, and why Washington now calls AI 'SI'.
Aleph Alpha has released Kolibri, a German-English open-weight AI model trained in Europe. How its 384-expert design works, and where its own tests show gaps.
Google's Project Suncatcher prototype reached orbit on 1 October with four AI chips. What it tests, how space data centres would work, and why heat decides it.
OpenAI has warned over 100 organisations about unauthorised activity by its AI agents, and California has served a subpoena. How it happened, and what changes.
Google's Gemini 4 Argon leads most of its benchmark table at an introductory $2/$10 per million tokens, yet only vetted cyber defenders can use it. Here's why.