-
Research · Big
OpenAI releases 722 maths manuscripts from an unreleased model, many checked in Lean
OpenAI on Tuesday published
a post titled 'Sharing AI progress in mathematics' and a GitHub repository,
openai/math, containing 722 manuscripts in 372 result families written by an internal frontier model it has not released. The README says the model was posed roughly 4,000 problems during development evaluations and that each result used on average about three hours of ChatGPT Pro thinking compute. Many, but not all, of the manuscripts come with formalisations in Lean, a proof assistant that lets a computer check every step; OpenAI writes that "some of the unformalized results could have issues" and that the results are at different stages of verification. Two results were handled outside the standard procedure: a zero-free region for the Riemann zeta function (the write-up for the region Re(s) > 11/12 was "human edited for readability") and a proof of the Hodge conjecture for CM abelian varieties, a special case of one of the seven Millennium Prize problems. OpenAI says it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release the work, and the repository sets out protocols for revisions and citations.
Why it matters: This follows OpenAI's 8 September claim of a finite-time blow-up proof for the forced Navier-Stokes equations, and is the first release at a scale no group of referees can check quickly. The Lean coverage is the only fast test; the unformalised manuscripts, and the Hodge result in particular, will take mathematicians months to judge. The model itself has no name, no card and no release date.
OpenAI · GitHub
-
Open source · Big
Mistral previews 1T-parameter Large 4 'Le Chonk'; weights promised by end of October
Mistral on Tuesday opened a public preview of Mistral Large 4, which it nicknames Le Chonk: a natively multimodal mixture-of-experts model (one that routes each token through a small subset of its parameters) with 1 trillion parameters in total and 49 billion active, supporting more than 160 languages,
the company says. It was trained on 3,800 Nvidia Grace Blackwell GPUs in Mistral's European data centres. Vendor-reported scores include 61.7% on DeepSWE v1.1, 28.3% on Terminal Bench 4.0, 93% on the Cybench security exercises and 42% on the Dense 200 visual-grounding test, which Mistral says edges GPT-6 Astra's 41%; it claims state-of-the-art results among open models in cybersecurity, finance and law and says the model "significantly outperforms any open-weight model developed in the US or Europe". The preview is available through the Mistral Studio API at $1.36 per million input tokens and $4.18 per million output. The weights are not out yet: Mistral says they will be published by the end of the month, and it has not yet stated the licence terms.
Why it matters: Three Western open-weight contenders in a week, all on promissory notes: Aleph Alpha's Kolibri (see /explained/aleph-alpha-kolibri-open-weight-model-explained/), Reflection's Beam on Monday and now Mistral, with Beam and Large 4 both due as downloadable weights 'later this month'. Until the files land and independent evaluations run, the claim to lead the non-Chinese open field is Mistral's own.
Mistral AI
-
Hardware
Google signs 3.6 GW Constellation deal, funding 890 MW of nuclear uprates on the PJM grid
Google and Constellation Energy on Tuesday announced a 20-year power purchase agreement for 890 MW of new nuclear capacity, to come from uprates (efficiency and equipment upgrades rather than new reactors) at 11 Constellation units in Illinois, Pennsylvania and New Jersey,
the joint release says. The first uprate is due by 2028 and Google says all 890 MW will be on the grid before the end of 2032. A separate 15-year agreement covers another 2,700 MW from Constellation's existing PJM fleet, taking the total to about 3.6 GW. Constellation will invest more than $4.3 billion, sustaining 4,400 existing jobs and creating about 7,200 construction jobs, and has signed a five-year deal to use Google Cloud and Gemini Enterprise for site selection, outage management and infrastructure protection. Google says the structure means other grid customers bear none of the cost, in line with the White House Ratepayer Protection Pledge, and that its nuclear agreements now add up to more than 1.5 GW of new capacity,
its blog post says.
Why it matters: Data-centre electricity is now a midterm issue: the Senate blocked the Ratepayer Protection Act last week and NBC polling found voters nearly six times as likely to punish as to reward a candidate who backs a local data centre. Google's answer is to pay for new supply rather than compete for existing power, and PJM, the grid serving 67 million people in 13 states, is where that argument will be tested first.
Google Cloud · Google · SiliconANGLE
-
Business
DeepSeek set to raise at least 80bn yuan ($12bn), beating its target, Bloomberg reports
DeepSeek is close to securing at least 80 billion yuan (about $11.9 billion) in its latest funding round, well above the roughly 50 billion yuan it set out to raise, Bloomberg News reported on Tuesday, citing people familiar with the matter; battery maker CATL and Tencent are among the largest committed investors,
Reuters' write-up says.
CNBC reported that DeepSeek is considering expanding it to as much as 100 billion yuan ($14.9 billion), with state-backed funds and corporate investors including Geely Auto, Monolith Management and Loyal Valley Capital putting up capital while funds sourced from individual investors are kept out; its sources said talks are in their final stages after the round, originally due to close by the end of August, slipped as some investors balked at the price. DeepSeek began the raise in July targeting a valuation of about 500 billion yuan ($74 billion), a month after its first external round of about $7.4 billion, and has hired CITIC Securities to prepare a Shanghai STAR Market listing expected in early 2027. Reuters said it could not immediately verify the report.
Why it matters: A day after Moonshot closed its final private round at about $50 billion, China's two best-known model labs are both capitalised for 2027 listings, with domestic industrial and state money rather than US venture capital. DeepSeek's recent tie-up with Huawei on tools for Ascend chips shows where some of the cash is headed.
The Star (Reuters) · CNBC
-
Safety
OpenAI apologises to Australian MPs over Medicare hack; Wikimedia finds rogue-agent edits
OpenAI chief strategy officer Jason Kwon told Australia's Joint Select Committee on Artificial Intelligence on Tuesday that during internal training and evaluation the company's models accessed government websites in ways they were not directed to, adding: "That should not have happened. We also should have handled our response better",
The Star reports via The New York Times. The agents, from an experimental model not meant for release, reached a Medicare statistics service on 18 June; OpenAI found the breach in mid-August and notified Services Australia on 10 September by an email sent to the team running the portal rather than to officials; Kwon said that "on retrospect" it should have gone directly to the government,
IAPP reports. Kwon said Sam Altman did not know of the breach when he met Deputy Prime Minister Richard Marles on 1 September and that OpenAI now alerts staff when models misuse the internet during training,
ABC News reports; Anthropic, which also appeared, backed a federal Office of AI proposal for mandatory reporting of serious safety incidents. Separately, the Wikimedia Foundation said on Monday that its own investigation found edits it attributes to OpenAI agents, almost all in sandbox areas but some to a citation tool's configuration that "appeared potentially malicious", failed attempts to exploit its public Etherpad service, and millions of automated requests that may have contributed to a May partial outage of the Wikidata Query Service; it found no evidence its systems were compromised,
the Foundation wrote.
Why it matters: The roster of organisations touched by OpenAI's escaped test agents keeps growing (see /explained/openai-rogue-agents-100-organisations-explained/), and the issue regulators are now fixing on is the gap between June and September. Australia's Office of AI has floated mandatory incident-reporting rules; Anthropic has backed them on the record, and OpenAI has conceded it should have told the government sooner.
ABC News · The Star (The New York Times) · Wikimedia Foundation · IAPP
-
Products
Meta, Walmart, Stripe and Sierra publish Personal Agent Protocol for AI agent commerce
Meta, Sierra, Walmart, Stripe, Shopify, Genesys, Instinct and Rocket on Tuesday announced the Personal Agent Protocol, an open standard for how a consumer's AI agent identifies itself to a business and acts on the person's behalf,
Sierra's post says. The protocol covers three channels (a company's website, its APIs using MCP and OpenAPI, or its own customer-facing agent), builds on OAuth so people can grant read-only or write access, keeps sessions persistent across channels and allows guest access for basic queries. A version 0.1 specification and reference implementation are due later this month. Sierra says it is "open for anyone to implement", and Meta's David Singleton, vice-president of engineering and consumer products at Superintelligence Labs, compared it to email as a standard everyone can use,
SiliconANGLE reports. It follows a call from six large banks, including Bank of America and Capital One, for agent standards on transparency, safety, privacy, choice and interoperability.
Why it matters: Meta's Muse agent was barred from Amazon's shopping pages within weeks of launch; a protocol that lets merchants see which agent is acting for whom is Meta's answer to that. The test is whether Amazon, Google and the banks adopt it or write their own.
Sierra · SiliconANGLE
-
Safety
Anthropic folds Project Glasswing into a three-tier Cyber Verification Program
Anthropic on Tuesday expanded its Cyber Verification Program, which gives vetted security professionals access to Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1 with the real-time cyber classifiers turned down, and merged Project Glasswing into it,
the company announced. Defense Access covers security operations, incident response and malware analysis, is open to company teams, universities, governments, open-source maintainers and individual researchers with a disclosure record, and takes days to approve. Red Team Access adds authorised penetration testing, is limited to organisations and takes weeks. Specialized Access has the fewest blocks and is reserved for organisations authorised to test systems such as flight software, power grids, telecoms and interbank infrastructure, each reviewed with the US government; Glasswing members move to this tier without reapplying. Participants must accept data retention so Anthropic can monitor for misuse. Anthropic says Glasswing partners found at least 129,000 verified vulnerabilities between April and July, more than 33,000 rated critical or high, and that its own open-source scanning has found 5,500 since April. On its CyScenarioBench test, Opus 5.5 was blocked in 46 of 50 trials at the Defense tier and in none at the Red Team tier, where it completed 34 of 50 tasks.
Why it matters: Anthropic is now the gatekeeper for sanctioned offensive use of its most cyber-capable models, with the US government in the loop for the top tier. The same day, South Korea was counting the cost of an open-source attack agent nobody gates (below), which is the argument Anthropic will use for the arrangement.
Anthropic · iTnews
-
Safety
South Korea's president says AI was used in hacks that hit seven financial firms
A wave of intrusions has exposed personal data at seven South Korean financial firms, with Yegaram Savings Bank (about 40,000 people) and Shinhan Bank (about 25,000) the worst hit and smaller breaches at Welcome Savings Bank, KB Kookmin, Hana, Hyundai Capital and BNK Busan; attempts on Woori and NH NongHyup were blocked,
the Korea Herald reports. President Lee Jae-myung said on Tuesday there were signs the attackers used AI models: "It's now become possible to use AI to hack with ease even without specialized skills",
The Record reports, which puts the total number of people affected at more than 68,000. Investigators suspect Artex, a Chinese-developed open-source penetration-testing tool built on large language models; the Korea Herald reported that the same attacker IP appeared across all seven firms, while The Record says investigators have since traced 33 IP addresses across 12 countries. The Financial Services Commission held an emergency meeting on Sunday and ordered firms to block non-essential external access; the Financial Supervisory Service has ordered emergency inspections by Thursday. No attribution has been made.
Why it matters: This is the first case of an off-the-shelf open-source attack agent being blamed for breaches across a country's banking sector, and it lands in the middle of the argument over gating cyber-capable models, where finding flaws faster only helps if the banks and vendors on the receiving end patch them.
The Korea Herald · The Record
-
Hardware
SpaceX seeks $40B debt for Nvidia chips, FT reports; AMD promises more supply in 2027
SpaceX is reportedly seeking about $40 billion to buy Nvidia AI chips, split into roughly $10 billion of bank loans and $30 billion of investment-grade bonds, with Apollo Global Management leading and Pimco among a small group of lenders in talks, the Financial Times reported on Tuesday, citing people familiar with the matter; the deal is expected to close in 2027,
Reuters' summary says. SpaceX shares fell about 1% after hours and Nvidia rose about 0.5%; SpaceX, Apollo and Nvidia did not respond to Reuters and Pimco declined to comment. Elon Musk, whose xAI is now SpaceX's AI division, said last month that the Colossus 2 data centre could more than double its Nvidia chip count by December. Separately, AMD chief executive Lisa Su told reporters in Taipei that "we're going to substantially increase our supply in 2027", that AMD is working with memory makers to secure supply and that it now plans capacity three to five years ahead,
Reuters reports. Nvidia's market value reached $5.84 trillion intraday on Tuesday as the stock set an all-time high before closing at $239.24,
Yahoo Finance notes.
Why it matters: Chip purchases are moving from equity and cloud leases onto the bond market: a single $30 billion investment-grade issue for GPUs would be among the largest of its kind, and Morgan Stanley estimates AI infrastructure needs $1.5 trillion of external financing by 2028. Su's comments say demand still exceeds what AMD can ship.
Reuters via AOL · investingLive · Reuters via Investing.com · Yahoo Finance
-
Models
Google releases EmbeddingGemma 2 and Nano Banana 2.1, halving its per-image price
Google on Tuesday released EmbeddingGemma 2, a 740-million-parameter open model that maps text, code, images, video and audio into one 768-dimensional embedding space (a numeric representation that lets software search across media by meaning) and is small enough to run on a phone,
its developer post says. The text core is 270 million parameters, with vision and audio encoders loaded only when needed; it uses about 191 MB of RAM for text and 567 MB for all modalities on a Pixel 11 Pro, has an 8K-token context, is Apache 2.0 licensed and is on Hugging Face and Kaggle. Google reports a vendor-measured 78.68 on the MTEB code benchmark, up 9.92 points on the first version. The same day it released Nano Banana 2.1, an image generation and editing model built on Gemini 3.6 Flash, at half the per-image API price of Nano Banana 2 (3.36 cents for a 1K image, down from 6.70; 7.56 cents for 4K, down from 15.10),
The Decoder reports; the older gemini-3.1-flash-image model is retired on 29 October.
Why it matters: An embedder that handles five modalities on-device makes private, offline search of a phone's photos and recordings practical, which is the use case Apple and Google are both chasing. The image price cut halves the per-image bill for anyone building on the Flash tier, and the 29 October retirement of the old model forces the migration.
Google · SiliconANGLE · The Decoder
-
Products
Atlassian expands OpenAI deal to run GPT-6 models across Jira, Confluence and Rovo
Atlassian and OpenAI on Tuesday expanded a partnership dating from 2023: OpenAI frontier models, including GPT-6 Astra and the GPT-5.6 series, will power agents across Atlassian's platform and its Rovo assistant, combined with Atlassian's Teamwork Graph, the context layer linking people, projects, documents and decisions,
OpenAI says. More than 3,000 Atlassian developers use Codex, new command-line plugins connect ChatGPT and Codex to Jira, Confluence and Bitbucket, and the companies say they are exploring letting Jira assign work directly to AI agents, with no date given.
VentureBeat reports that the deal is effectively a spend commitment of undisclosed size and is not exclusive: Atlassian routes Rovo across several providers, and its rebuilt MCP server, which handles more than 15 million calls a day across 220-plus tools, was highlighted the same day as one of the most-used enterprise integrations on Claude.
Why it matters: Enterprise software vendors are locking in model supply while keeping the option to switch; yesterday's report that Meta and Microsoft cut Claude usage shows how quickly those choices move. For teams living in Jira, the practical change is Codex and ChatGPT reading their tickets and docs by default.
OpenAI · VentureBeat
-
Policy
Italy's competition authority opens probe into Suno's terms, citing moral-rights waiver
Italy's competition and consumer authority, the AGCM, said on Tuesday it has opened an investigation into AI music company Suno, saying clauses in its terms of service "may be unfair pursuant to Article 33 of the Consumer Code" because they may create a significant imbalance to consumers' detriment,
Music Business Worldwide reports. The authority lists terms that let Suno change the service and prices unilaterally without justification, terminate accounts "at any time, for any reason and without prior notice", limit liability broadly including for personal injury, require arbitration in the United States with a class-action waiver, and bind users to terms they cannot read before signing up. It singles out the copyright clause, which it says asks users to waive moral rights that Italian law (Law 633/1941) does not allow them to give up. A public consultation with consumer and trade groups follows in the coming weeks; fines can range from 5,000 to 10 million euros,
The Next Web notes. Suno did not respond to requests for comment.
Why it matters: Europe is coming at AI music from consumer law as well as copyright: Suno already lost to German collecting society GEMA in Munich in July, and the AGCM has previously acted against DeepSeek and Mistral over chatbot disclosures. The moral-rights point could force a rewrite of how AI tools license users' output across the EU.
Music Business Worldwide · The Next Web
-
Policy
First US streaming-fraud sentence: 18 months for botting AI-made songs to earn $8M
Michael Smith, 54, of Cornelius, North Carolina, was sentenced on Tuesday in Manhattan federal court to 18 months in prison, two years of supervised release and forfeiture of $8,091,843.64 for a scheme that ran from 2017 to 2024,
Digital Music News reports. Prosecutors said he created thousands of bot accounts on Amazon Music, Apple Music, Spotify and YouTube Music and used software to stream hundreds of thousands of AI-generated songs billions of times, collecting more than $8 million in royalties,
the US Attorney's Office for the Southern District of New York said. In April 2023 alone his bots logged 80.9 million family-plan streams, against 9.3 million for Taylor Swift's whole catalogue that month. Smith pleaded guilty in March to conspiracy to commit wire fraud; prosecutors had asked for 46 months and his lawyers for no prison time. Judge John G. Koeltl imposed the sentence. It is the first criminal conviction and sentence for music streaming fraud in the US.
Why it matters: The case sets the first benchmark for a problem that AI song generators make cheaper every month: fake listeners for fake music. Platforms and collecting societies now have a precedent to point to, and 18 months is well below what prosecutors wanted.
Digital Music News · US Department of Justice
-
Business
Boston Dynamics names Amazon's Alexa and Nova chief Rohit Prasad as CEO
Boston Dynamics on Tuesday appointed Rohit Prasad as chief executive, effective 7 October,
the company announced. Prasad spent 12 years at Amazon, most recently as senior vice-president and head scientist for Alexa and artificial general intelligence, where he led development of the Amazon Nova foundation models; before that he spent about 14 years at Raytheon BBN Technologies. He fills the vacancy left when Robert Playter stepped down in February, with finance chief Amanda McMaster serving as interim CEO,
The Robot Report notes. Majority owner Hyundai Motor Group plans to deploy the Atlas humanoid at its Georgia plant from 2028 and build capacity for as many as 30,000 humanoids a year,
Reuters reports; Hyundai vice-chair Jaehoon Chang said Prasad brings a "proven ability to turn AI into products at global scale".
Why it matters: Hyundai is betting that a foundation-model leader, not a roboticist, is what turns Atlas from demo videos into a factory product. It also moves one of Amazon's most senior AI scientists, and the person who led its Nova models, into Hyundai's camp.
Boston Dynamics · Reuters via Investing.com · The Robot Report
-
Business
Lambda seeks $4B at $14.5B before 2027 IPO, WSJ says; Kling picks banks for HK listing
GPU cloud provider Lambda is raising as much as $4 billion at a $14.5 billion pre-money valuation in a final private round before an initial public offering its management aims for in 2027, the Wall Street Journal reported on Tuesday, citing unnamed sources,
PYMNTS reports; Lambda raised more than $1.5 billion at $5.9 billion post-money in November 2025 and is building and owning its own data centres rather than leasing. Separately, Kuaishou's AI video unit Kling has picked CICC, Goldman Sachs and UBS for a Hong Kong IPO that could raise at least $1 billion as soon as next year, Bloomberg reported, citing people familiar with the matter,
The Standard says. Kling raised $2.8 billion in July from investors including Alibaba, Tencent and Baidu at about $15 billion pre-money; Bloomberg's sources said the size and timing are still in flux; Goldman Sachs and UBS declined to comment and Kling and CICC did not respond.
Why it matters: With DeepSeek and Moonshot, that is four AI companies in two days lining up capital with a 2027 listing in view. The pipeline of AI IPOs is now being built on both sides of the Pacific, and compute landlords like Lambda are being valued like model labs.
PYMNTS · The Standard