Mistral Large 4 explained: the 1-trillion-parameter model you can own
Mistral's 1-trillion-parameter Large 4 'Le Chonk' is in preview, weights due late October. How its sparse design works, what it scores, and what is unproven.
In 60 seconds
- On 6 October Mistral opened a public preview of Mistral Large 4, nicknamed 'Le Chonk': a mixture-of-experts model with about 1.05 trillion parameters, roughly 50 billion of them active per token, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Europe. Weights are promised by the end of October (27 October, per VentureBeat) under a licence not yet published.
- Each token is routed through a small subset of the model's expert networks, so the compute per token is that of a 50-billion-parameter model while the memory footprint is that of a trillion-parameter one.
- Independent testing by Artificial Analysis places it first among open models built outside China but behind seven Chinese open models. Its strongest result, 82% on a reproduce-and-patch vulnerability test, is independently measured too; the 'near zero' Mistral quotes for Claude Opus 5.5 and GPT-6 Astra reflects refusals, not failed attempts.
Mistral has shown the model that a cat meme predicted. On 6 October the Paris company opened a public preview of Mistral Large 4, a roughly one-trillion-parameter model it calls "Le Chonk", trained from scratch in its own European data centres, with downloadable weights promised by the end of October (27 October, according to VentureBeat). The benchmarks are a mixed picture, but the argument Mistral is making is sharper than the scores: that in a world where a US provider can withdraw or restrict a model, a copy you hold yourself is the only one you can plan around.
For: EveryoneThe plain-English version
Think of a large language model as a very big reference library with a librarian. When you ask a question, the librarian does not read every book. They walk to the two or three shelves that matter and read those. Mistral Large 4 is built this way. It holds about 1.05 trillion "parameters", the numbers that encode what it has learned, which is the whole library. But for each word it produces it consults only about 50 billion of them, the few relevant shelves. That design is called a "mixture of experts", and it is why a model this large can answer at a speed Artificial Analysis measured at about 116 words-worth of tokens per second. The catch is that the library still has to exist. You cannot run the model on a laptop; you need a rack of server chips holding roughly a terabyte of memory.
Mistral's announcement, published on its site, makes three claims. First, that the model was built entirely in Europe, on 3,800 NVIDIA Grace Blackwell chips in Mistral's own data centres, trained from scratch rather than copied from someone else's model. Second, that it is the strongest model you will be able to download that was built outside China. Third, that it is unusually good at computer security work, like finding a flaw in software and writing a fix.
All three claims have outside support, with caveats. An independent evaluation firm, Artificial Analysis, scores it at 38 on its general intelligence index, which, as the analysis site mrkt30 tallied, is above every other open model from the US or Europe but below seven Chinese ones. The same firm measured the security result too: 81.7% on its reproduce-and-patch test, the top score on its board. But the third claim comes with an asterisk we unpack below: the closed American models Mistral's chart shows at "near zero" did not fail that test, they refused to take it.
Why does any of this matter to a non-specialist? Because "open-weight" means a hospital, a bank or a government can download the model and run it on its own machines, with nobody able to switch it off or read its data. Mistral's chief scientist Guillaume Lample put the pitch bluntly, as Implicator reported: "If you use a closed model, there is no guarantee it will still be there tomorrow." The company is betting that for a growing number of buyers, especially in Europe, that guarantee matters more than being the single best model.
And the name? In June a joke model called "Le Chaton Fat", the fat kitten, went viral with made-up benchmarks and a claimed 30 trillion parameters. Mistral's chief executive Arthur Mensch played along, and the real model inherited the cat.
For: CuriousHow it actually works
To see what Mistral built, it helps to separate four things that get blurred together in the coverage: the architecture, the training, the "open" part, and the cybersecurity positioning.
The architecture. Mistral's model page describes Large 4 as a "granular Mixture-of-Experts" model with 1.05 trillion total parameters and 52 billion active, plus a separate 1.6-billion-parameter vision encoder so it can read images, charts, PDFs and technical drawings. (VentureBeat, Artificial Analysis and most other coverage give 49 billion active; the discrepancy is unexplained, and we use "about 50 billion".) In a dense model every parameter touches every token. In a mixture-of-experts model, each layer's feed-forward block is split into many smaller expert networks, and a small router network decides, token by token, which few experts to use. "Granular" signals many small experts rather than a handful of large ones, the same direction Aleph Alpha took with Kolibri. The payoff is a ratio: roughly 5% of the parameters do the work for any given token, so the compute bill scales with 50 billion while the knowledge capacity scales with a trillion.
The training. Mistral says it trained Large 4 from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its European data centres; VentureBeat and Artificial Analysis round that to 4,000, and VentureBeat reports the run took about two months. The previous model, Large 3, used 3,000 older H200 GPUs and had 675 billion total and 41 billion active parameters. After pre-training comes post-training: supervised fine-tuning and then reinforcement learning, where the model attempts tasks, is scored, and is nudged toward higher-scoring behaviour. Mistral gives an unusual operational detail here: on about 3,000 GPUs its RL pipeline produces around 33 billion tokens a day, of which about 16 billion are "trainable completion tokens" after filtering. It also says the RL run "is still in progress" and "shows no signs of saturation", which is a polite way of saying the preview is not the final model.
The "open" part. Mistral's post promises the weights "by the end of the month"; VentureBeat reports 27 October, following a roughly three-week testing period with developers, security firms and government authorities. The weights are expected, per VentureBeat, under a custom Mistral licence rather than the permissive Apache 2.0 that Large 3 used. Until then the model is available only through Mistral's API, at a list price of $1.36 per million input tokens and $4.18 per million output tokens, currently shown at half price as a launch "sale" on the docs page.
The cybersecurity positioning. This is the part that makes Large 4 a story rather than a release. Mistral's thesis is that security teams need a model that will do offensive-looking things for defensive reasons: reproduce an exploit to confirm a vulnerability is real, write a patch, analyse malware, write detection rules. Closed providers often refuse such requests. Mistral's chart shows Large 4 scoring 82% on a reproduce-and-patch test (Artificial Analysis's CyberGym-E2E, where the firm's own run gives 81.7%) and, it says, Claude Opus 5.5 and GPT-6 Astra score near zero because they refuse. To release a model with those abilities responsibly, Mistral says it is red-teaming a version "with reduced moderation and expanded cyber capabilities" with "cybersecurity leaders, vetted partners, and state authorities" before the weights go out.
Put the four together and the logic is coherent: a large but efficient model, trained on European hardware, that you can eventually download, aimed at buyers for whom dependency on a US API is the problem. Whether the execution matches the logic is what the numbers are for.
For: PractitionerThe deep dive
Sparse scaling: what 1.05T / ~50B buys you
For a mixture-of-experts transformer the per-token training cost is approximately
floating-point operations, where counts only the parameters touched per token (the factor of 6 covers the forward and backward passes). Mistral has not disclosed its token count, but the announced hardware lets us bound it. Taking 3,800 GB200-class GPUs at roughly 2.25 petaFLOPS dense BF16 each, a two-month run and a 40% utilisation rate gives about FLOPs. Dividing by yields on the order of 55 trillion training tokens. That is our back-of-envelope estimate, not a Mistral figure, and every input to it is approximate; the point is the shape. A dense trillion-parameter model on the same budget would see roughly twenty times fewer tokens. Sparsity is what makes a trillion parameters trainable on a few thousand chips in a few weeks.
Inference inverts the trade. Compute per generated token is that of a 52-billion-parameter model, which is why Artificial Analysis measures 116 tokens per second and a 1.46-second time to first token, both better than the median of comparable models. But every expert must be resident, because the router may call any of them for the next token. At 8-bit precision 1.05 trillion parameters is about 1.05 terabytes before activations and the key-value cache for a long context. A single eight-GPU node of 141 GB H200s (1.13 TB) is marginal; realistically self-hosters are looking at Grace Blackwell-class nodes or multi-node tensor and expert parallelism. The Developers Digest analysis is right that this is "not the same operational object as a 27B local coding model": serving topology, memory bandwidth and quantisation quality will decide the real cost.
Context length is another place where sources disagree. Mistral's docs say 1 million tokens; Artificial Analysis lists 524k, the size it tested at. Treat the larger figure as an API ceiling that has not been independently characterised.
The RL pipeline numbers
Mistral's disclosure that its RL run generates about 33 billion tokens a day on roughly 3,000 GPUs, of which about 16 billion are trainable completion tokens, is a rare look at post-training throughput. The halving reflects the filtering and masking that reinforcement learning on agentic tasks requires: prompts, tool outputs and environment observations are generated but not trained on, and failed or malformed rollouts are discarded. At that rate a month of RL contributes under half a trillion trainable tokens, roughly 1% of our pre-training estimate, which is consistent with the general finding that post-training is compute-light but data-expensive. Mistral says the same environment is what it sells to customers as Mistral Forge, so the disclosure doubles as a product demo.
Benchmarks: who measured what
Mistral's launch post leans on third-party harnesses for coding and on its own runs elsewhere. The table separates them.
| Benchmark | Large 4 | Comparison | Measured by |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 38 (rank 64 of 225) | MiMo-V2.6-Pro 46, GLM-5.3 45, Kimi K3 44, DeepSeek V4.1 Flash 39; Large 3 scored 9 | Artificial Analysis, independent (AA, mrkt30) |
| Cost per Intelligence Index task | $1.13 | Used 200M output tokens vs 81M median | Artificial Analysis |
| DeepSWE v1.1 | 61.7% | Mistral's chart: GLM-5.3 61%, DeepSeek V4 Pro 57%; live leaderboard best configs: GLM-5.3 and Kimi K3 ~69%, GPT-6 Astra, Gemini 3.8 Flash and Opus 5 ~74% | Mistral's chart; leaderboard figures via VentureBeat |
| Terminal-Bench 4 | 28.3% (Mistral); 22.7% (Vals AI's own run) | GLM-5.3 about 42% in third-party compilations of Artificial Analysis data | Mixed |
| SWE-Atlas-QnA | 59.4% | Not stated | Vendor-cited |
| Surge AI blind human coding rating (1 to 5) | 3.74, 2nd of 5 | Claude Opus 5 4.22, GLM-5.3 3.60, Kimi K3 3.59 | Surge AI evaluation, reported by Mistral |
| AutomationBench (657 workflows) | 59.9% | Decrypt's chart: Claude Sonnet 5.5 71.8, Opus 5.5 69.5, Gemini 4 Argon 77.5 | Vendor-reported |
| Harvey Legal Agent | 15% (Mistral) / 15.8%, 6th of 76 (Vals) | Kimi K3 12.92%, MiMo V2.6 Pro 10.83%, GLM-5.3 8.33% | Vals AI |
| Finance Agent v2 | 54.7%, 22nd of 76 | GPT-6 Astra 53.5, Claude Opus 5.5 58.6 | Vals AI; comparison chart via Decrypt |
| CyberGym-E2E reproduce-and-patch (131 tasks) | 82% (Mistral); 81.7% (AA) | AA board: MiMo-V2.6-Pro 78.6%, GPT-6 Luna (Max) 77.9%; Mistral's chart puts Claude Opus 5.5 and GPT-6 Astra "near zero" due to refusals | Artificial Analysis, independent |
| Cybench (40 challenges) | 93% | Not stated | Vendor-reported |
| Artificial Analysis Cyber Index | 50 | MiMo-V2.6-Pro 56; GLM-5.3-Flash 50; "top five globally" is Mistral's framing | Artificial Analysis, independent; score via mrkt30 |
| Lakera B3 prompt-injection resistance | 93.3% of attacks resisted | Not stated | Vendor-reported |
| Dense 200 visual grounding | 42% | GPT-6 Astra 41% | Vendor-reported |
Three patterns emerge. First, on the broad independent index Large 4 is a clear generational jump for Mistral (9 to 38) and the best open model from outside China, but eighth among open models overall. Second, coding is the weak spot: on Terminal-Bench 4 it is at roughly two-thirds of GLM-5.3, and on the live DeepSWE board the models Mistral compares against score higher in their best configurations than in Mistral's chart. VentureBeat's conclusion, that the result "does not establish an outright coding lead across every available model-and-agent configuration", is fair. Third, the cyber picture is split: the reproduce-and-patch score is independently confirmed and genuinely leads Artificial Analysis's board, but the composite Cyber Index puts Large 4 at 50, behind MiMo-V2.6-Pro's 56, and the Cybench and agentic numbers are still Mistral's own.
The refusal comparison, examined
The 82% versus "near zero" claim deserves a closer look because it is the launch's most quoted number. The task is Artificial Analysis's CyberGym-E2E-AA: 131 memory-safety vulnerabilities, one per open-source project, each with a 90-minute limit in a sandboxed agent harness. The model must reproduce the vulnerability and then patch it. Artificial Analysis's own run gives Large 4 Preview 81.7% pass@1, ahead of MiMo-V2.6-Pro at 78.6% and GPT-6 Luna (Max) at 77.9%, so the score itself is solid. The "near zero" half of the chart is different. Reproducing a vulnerability means writing a working proof-of-concept exploit. Anthropic's and OpenAI's flagship models decline to do that under their usage policies, so they score at the floor, and Artificial Analysis's public board does not list them. The comparison therefore measures a policy choice, and the honest framing is Mistral's own argument stated plainly: it has chosen a different refusal boundary and believes defenders need it. That is a legitimate position with real trade-offs, and it is the same tension Google wrestled with in its gated Gemini 4 Argon release. It is not evidence of a capability gap.
Mistral's safety numbers point the other way: it reports the highest refusal rate for harmful cyber requests among open models it compared on JailbreakBench, StrongREJECT and AgentHarm, a KORA score of 1.691 out of 2, and 93.3% resistance on Lakera's B3 injection benchmark. The public API, in other words, is tuned to refuse malicious requests while answering defensive ones. The reduced-moderation version exists only for vetted partners during the preview. What happens to that boundary once the weights are public, and anyone can fine-tune it away, is the open question an open-weight release cannot answer.
Compared with the closest prior work
The nearest comparison is Moonshot's Kimi K3, a 2.8-trillion-parameter open-weight mixture-of-experts model (104 billion active) released in July, and Z.ai's GLM-5.3. Both outscore Large 4 on Artificial Analysis's index and on coding. Large 4's distinguishing features are provenance and price-performance: trained in Europe on European-owned hardware, 160-plus languages including every official EU language, and output pricing that mrkt30 notes is roughly level with GLM-5.3 at list but about five times MiMo-V2.6-Pro's $0.87 per million output tokens. Against Mistral's own Large 3 the jump is unambiguous.
For: EveryoneWhy it matters
For everyday users, the direct effect is small and indirect. Large 4 will power Mistral's assistant products and it is multilingual in a way few frontier models are, with every official EU language covered. The larger effect is that a credible non-US, non-Chinese model exists at this scale, which shapes what governments and companies in Europe will deploy in the services people use.
For developers and builders, the preview is usable now through Mistral's API at a price well below the closed frontier: Decrypt's comparison puts it at roughly a third of Claude Opus 5.5 on input and a fifth on output. Early hands-on reports are mixed. On Hacker News, Simon Willison noted that the model "only supports reasoning 'none' or reasoning 'high'" and that the setting "didn't seem to make any real difference" in his tests, and Stork relayed a test by Matthew Berman in which "excessive reasoning output quickly exhausted the context window" in an agentic coding harness. Verbosity is also what Artificial Analysis measured. Developers should budget by tokens per task, not per million.
For companies, especially regulated ones, the promised weights are the product. A downloadable trillion-parameter model with a vendor willing to support on-premise deployment is a procurement option that did not exist from a Western lab at this quality level. The cost is hardware: a terabyte-class GPU node is a capital decision, not a credit-card one.
For the field, Large 4 is a data point in two running arguments. One is whether a few thousand GPUs can still produce a near-frontier model; Implicator notes Jensen Huang's figure of about 100,000 GPUs for GPT-6 Astra, and while the numbers are not comparable in any precise way, the order-of-magnitude gap is real and Mistral landed within a few index points of several models trained with far more. The other is the refusal boundary for security work. If a well-funded European lab ships open weights that will write exploits for defenders, the US labs' stricter policies become a competitive variable rather than an industry norm.
For: CriticalWhat to be skeptical of
The weights are not out. Every protection against lock-in that Mistral advertises depends on a release three weeks away under a licence nobody has read. The docs call the model "open-weight" but name no licence; VentureBeat reports it will be a custom one. Mistral's own precedent is uneven: Large 3 was Apache 2.0, while mrkt30 notes Medium 3.5's modified MIT licence excludes companies above $20 million in monthly revenue. If Large 4's licence carries similar thresholds, the sovereignty pitch narrows to those who can negotiate.
Several headline claims are still vendor-reported. The 93% Cybench result, the AutomationBench and SWE-Atlas figures and the "top five" Cyber Index framing come from Mistral. Artificial Analysis's composite Cyber Index gives Large 4 a 50, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro, so "top five globally" depends on which models you count. Mistral acknowledges that architecture details, additional benchmarks and the post-training methodology will come with the weights.
Chart presentation drew criticism. One Hacker News commenter, oh_no, pointed out that Mistral's chart showed GPT-5.6 Sol at 12.8 on a document benchmark where Artificial Analysis lists the Sol 5.6 Max configuration at 27, calling it "extremely weird to cherry pick" that model. VentureBeat could not verify several competitor figures in Mistral's visual-grounding and Finch charts against public sources.
The model is still changing. Mistral says the RL run is unfinished. That is good news for the final release but means today's numbers describe a checkpoint that will not ship.
The refusal argument has a cost. Open weights cannot be recalled, and a model tuned to write exploits for defenders writes them for attackers too. Pre-release red-teaming with vetted partners is a reasonable precaution, but it cannot bind anyone who downloads the weights later. Mistral has not published red-teaming findings or methodology.
Mistral's own people are measured. Pierre Stock, Mistral's vice-president of science, told Axios, as quoted by mrkt30, "We're not there yet on the frontier." The 49-versus-52-billion active parameter discrepancy and the 524k-versus-1M context discrepancy are small but suggest the preview's documentation was assembled quickly.
For: EveryoneWhat to watch next
- Late October: the weights and the licence. Mistral says "by the end of the month"; VentureBeat says 27 October. Check Mistral's Hugging Face page for a Large 4 checkpoint and read the licence terms, especially any revenue or use-case thresholds and whether redistribution of fine-tunes is permitted.
- Independent cyber scores. The 82% CyberGym result already has Artificial Analysis's confirmation; a Cybench replication by a party other than Mistral would settle whether the 93% holds, and the Cyber Index will show whether the released weights land in the top three open models, as Artificial Analysis expects.
- The final checkpoint's coding numbers. Mistral says RL is ongoing; the live DeepSWE and Terminal-Bench boards will show whether the gap to GLM-5.3 and Kimi K3 closes.
- Architecture disclosure. Expert count, routing scheme, attention layout and the training token count are all promised with the weights. They will let outsiders check the compute estimates above.
- The refusal boundary once weights are public. Watch for fine-tunes that strip the safety tuning, and for how Mistral's Forge customers and state partners actually use the reduced-moderation version.
- Europe's other open-weight bet. Aleph Alpha's Kolibri took the same sparse-expert route at about a thirteenth of the size. The two releases together are the clearest test yet of whether European labs can field models that regulated buyers will actually choose.
Check your understanding
Pick an answer — you'll see why right away.
1. Mistral Large 4 has about 1.05 trillion parameters but only about 50 billion are 'active' per token. What does that mean for an organisation that wants to run the weights itself when they are released in late October?
2. Mistral reports 82% on a test that asks a model to reproduce and then patch a real vulnerability, and says Claude Opus 5.5 and GPT-6 Astra score 'near zero'. What does that comparison actually show?
3. Artificial Analysis gives Mistral Large 4 an Intelligence Index score of 38 and notes it used about 200 million output tokens to complete the index, against a median of 81 million. Why does that second number matter for a buyer comparing it on price?
4. Why do several commentators say that, for now, 'open-weight' describes a promise rather than the product in Mistral's announcement?
Glossary
- Mixture of experts (MoE)
- A model design where each layer holds many small 'expert' networks and a router sends each token through only a few of them, so computation per token is far smaller than the total parameter count.
- Active parameters
- The number of parameters actually used to process one token in a mixture-of-experts model, as opposed to the total number stored in memory.
- Open-weight model
- A model whose trained parameters can be downloaded and run on your own hardware, under whatever licence the publisher chooses; not the same as open source, which also implies freely reusable training code and data.
- Reinforcement learning (RL) post-training
- A training phase after pre-training in which the model generates answers, is scored against a reward, and is updated to produce higher-scoring answers; used heavily for coding and agentic tasks.
- Grace Blackwell GPU
- NVIDIA's data-centre platform pairing a Grace CPU with Blackwell-generation GPUs, sold as the GB200 and GB300 systems; Mistral says it trained Large 4 on 3,800 of them.
- Cybench
- A benchmark of 40 capture-the-flag style security challenges drawn from real competitions, used to measure how well a model finds and exploits vulnerabilities.
- Prompt injection
- An attack where malicious instructions hidden in a document or web page try to hijack an AI agent; Lakera's B3 benchmark measures resistance to it.
- Artificial Analysis Intelligence Index
- An independent composite score from the evaluation firm Artificial Analysis that averages a model's results across a fixed set of reasoning, coding and knowledge benchmarks.
- CyberGym-E2E
- A test, run independently by Artificial Analysis, in which an agent must reproduce a real memory-safety vulnerability in an open-source project and then patch it, within 90 minutes per task across 131 tasks.
- Red-teaming
- Deliberately attacking a model to find harmful or unsafe behaviour before release, often with outside experts under agreement.
Questions people ask
Is Mistral Large 4 open source?
Not yet, and strictly speaking it will be open-weight rather than open source. Mistral says the weights will be downloadable by the end of October 2026 (VentureBeat reports 27 October), but as of 8 October no licence has been named and no checkpoint has been published. VentureBeat reports the licence will be a custom Mistral one, unlike Large 3's Apache 2.0.
How big is Mistral Large 4 and can I run it locally?
Mistral's docs give 1.05 trillion total parameters with 52 billion active per token; its launch post also says 52 billion, while VentureBeat, Artificial Analysis and most other coverage say 49 billion. Because all parameters must be in memory, you need roughly a terabyte of GPU memory at 8-bit precision. It is a multi-GPU server model, not a laptop one.
How does Mistral Large 4 compare to GPT-6, Claude Opus 5.5 and Chinese open models?
Independent testing by Artificial Analysis scores it 38 on its Intelligence Index, the highest for an open model built outside China, but below seven Chinese open models led by MiMo-V2.6-Pro at 46. Mistral's own charts compare it mainly to Chinese open models. Against closed frontier models it trails clearly on coding leaderboards such as DeepSWE.
What does Mistral Large 4 cost?
The list price is $1.36 per million input tokens and $4.18 per million output tokens. Mistral's docs currently show those halved to $0.68 and $2.09 as a 'sale price', which Artificial Analysis describes as a two-week launch discount. Artificial Analysis estimates about $1.13 per task on its index, partly because the model is more verbose than most.
Why is Mistral Large 4 called 'Le Chonk'?
In June 2026 a joke model called 'Le Chaton Fat' (the fat kitten) went viral on Reddit and X with fake benchmarks and claims of 30 trillion parameters. Mistral's chief executive Arthur Mensch played along, calling it 'le gros chaton', and the company adopted the cat as the model's unofficial mascot.
Why does Mistral emphasise cybersecurity for Large 4?
Mistral argues that closed-model providers refuse many legitimate security tasks, such as reproducing a vulnerability to patch it, and that a defender's model could be withdrawn mid-incident. It reports 82% on a reproduce-and-patch test (Artificial Analysis independently measured 81.7%) and 93% on Cybench (vendor-reported), and is red-teaming a reduced-moderation version with security firms and state authorities before releasing the weights.
Discussion
- Loading comments…
Sources
- Mistral Large 4 — Mistral AI · official announcement
- Mistral Large 4 model page — Mistral AI · docs
- Mistral Large 4 Preview: Intelligence, Performance & Price Analysis — Artificial Analysis · analysis
- Mistral debuts Large 4 'Le Chonk', a 1-trillion parameter text output model with high benchmarks planned for open weights release — VentureBeat · news
- Mistral Large 4 Preview Pitches Open Weights Against Lock-In — Implicator · analysis
- Mistral AI drops 'Le Chonk', a massive AI model named after a cat meme — Decrypt · news
- Mistral Large 4: Benchmarks vs China, Price and Licence — mrkt30 · analysis
- Mistral's Le Chonk has a huge catch — Stork · analysis
- Mistral Large 4 Is the Open-Weight Frontier Test — Developers Digest · analysis
- Mistral Large 4 (Hacker News discussion) — Hacker News · analysis
- Mistral Large 4 Debuts as 1T-Parameter Open-Weight Model Built for Cyber Defense — VKTR · news
- Mistral Large 4 (Le Chonk): Specs, Price, Benchmarks — CellCog · analysis
- CyberGym-E2E-AA leaderboard — Artificial Analysis · benchmark
- Mistral Large 4 on Vals AI — Vals AI · benchmark
- Mistral Large 4 on Artificial Analysis: Eighth Among Open Models — Trending Topics · news
How this was made: researched and written by an AI model (Claude) from the primary sources listed above, then checked claim-by-claim against those sources in a separate AI fact-check pass. Spotted an error? Email [email protected] and we correct it publicly. Our process.