Mistral Previews the 1T-Parameter Mistral Large 4: Open Weights by Month's End, 82% on a Reproduce-and-Patch Security Test, Still Well Behind on Terminal Agent Tasks

On October 6 Mistral AI released a public preview of Mistral Large 4: a mixture-of-experts model with 1 trillion total and 49 billion active parameters, combining instruct and reasoning modes, natively accepting text and image input, covering more than 160 languages, and trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in Mistral's own European datacenters. The preview API is live on Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens; Mistral says it will release the weights by the end of the month, with the license not yet stated. In Mistral's figures, it scores 82% on the Artificial Analysis Cyber Index test that asks a model to reproduce and patch a real open-source vulnerability, with Claude Opus 5.5 and GPT-6 Astra near zero because they refuse the task; it scores 61.7% on DeepSWE v1.1 and 28.3% on Terminal-Bench 4.

Model and price

On October 6 Mistral AI released a public preview of Mistral Large 4. It is a mixture-of-experts model with 1 trillion total parameters and 49 billion active per step, combining instruct and reasoning modes in one model, natively accepting text plus image input, and covering more than 160 languages including every official EU language. Mistral says it was trained from scratch on 3,800 Nvidia Grace Blackwell GPUs in its own datacenters in Europe.

The preview API can be tried today on Mistral Studio at $1.36 per million input tokens and $4.18 per million output tokens. The weights are due "by the end of the month," but the license hasn't been stated. Mistral VP of Science Pierre Stock told TechCrunch that until then the company will work with trusted partners to make sure the open weights can be used for defense rather than malicious attacks.

Scores: strong on security, weak on terminal tasks

Mistral's headline is cybersecurity: on the Artificial Analysis Cyber Index test that asks a model to reproduce a real open-source vulnerability and then patch it, it scores 82%, which Mistral calls the highest of any model, and it solves 93% of Cybench, a set of 40 security competition exercises. Mistral also notes that Claude Opus 5.5 and GPT-6 Astra score near zero on the same test because they refuse the task, so that result reflects differences in safety policy more than in capability.

On general agent work it doesn't lead. DeepSWE v1.1 is 61.7%, against 77.9% for Google's Gemini 4 Argon announced at the end of September and 74.2% for Opus 5.5; Terminal-Bench 4 is just 28.3%, where Argon and Opus 5.5 score 57.4% and 66.4%. In Mistral's own human coding evaluation it scores 3.74, below Claude Opus 5's 4.22 and slightly above Kimi K3's 3.59.

What it means

If the weights do ship this month under a permissive license, this will be the first trillion-parameter open-weight model from a European company, with real appeal for enterprises that need on-premises deployment and data kept inside the EU, and its price is well below US closed flagships. But by Mistral's own numbers, its strengths are cybersecurity, automation and multilingual work, while complex terminal and coding agent tasks remain a tier behind. All figures are self-reported, and the comparison basis for the security results deserves a careful read.

via: Mistral announcement, TechCrunch report