Performance and Benchmarks
Mistral launched its trillion-parameter Large 4 model, pushing the frontier of open-weight performance. Artificial Analysis, an independent AI benchmarking firm, ranked Mistral Large 4 as the best model available from outside the U.S. and China in its index. The mixture-of-experts model, trained on Nvidia Grace Blackwell GPUs, scored its personal best in the CyberGym-E2E-AA end-to-end cybersecurity measure, although several Chinese open models beat it overall.
Mistral Large 4 achieved a score of 38 on the Artificial Analysis Intelligence Index. The 1-trillion parameter model, featuring 49 billion active parameters, matches the performance of GPT-6 Luna and DeepSeek V4.1 Flash. It also achieved 50% on the Artificial Analysis Cyber Index, ranking level with GLM-5.3-Flash and outperforming models such as Kimi K3. Its strongest specific result is an 82% score on CyberGym-E2E-AA.
At least five Chinese open models scored higher on the intelligence index: DeepSeek V4.1 Flash at 39, GLM-5.3-Flash at 42, Kimi K3 at 44, GLM-5.3 at 45, and MiMo-V2.6-Pro at 46. While MiMo-V2.6-Pro is comparable in total size, GLM-5.3-Flash is significantly smaller at 320 billion total parameters and 18 billion active parameters. Mistral states that the model weights for Large 4 will be released by the end of October.
Cost and Efficiency
Mistral Large 4 costs $1.13 per benchmark task at list price, or $0.57 during its initial 50% discounted launch window. This compares to $0.13 for MiMo-V2.6-Pro and $0.25 for GLM-5.3-Flash. Despite the higher cost, the model secured an 81.7% score on CyberGym-E2E-AA.
The cybersecurity score is part of a broader index where the model scored 16% in DeepsecBench-AA and 51% in CWE-Bench-AA. Independent projections indicate that once open weights ship, the model will rank among the top three open-weight entries on the Cyber Index.
Enterprise Workloads
Mistral positions Large 4 as state-of-the-art among open models for enterprise workloads including cybersecurity, finance, and law. The model features a 1.6-billion-parameter vision encoder supporting document and image reasoning, accepting up to 100 images per API request.
Company executives highlighted that the model exceeds certain Chinese alternatives in cybersecurity capabilities, alongside benchmark scores of 61.7% on DeepSWE v1.1, 49.8% on the Coding Agent Index, and 59.9% on AutomationBench.
Large 4 outperforms OpenAI's closed GPT-6 Luna on the CyberGym-E2E-AA benchmark, though Luna matches the 38 score on the main intelligence index at a lower task cost of $0.07. Mistral noted that some frontier models decline specific tasks on safety grounds, scoring near zero.
The model also outperforms DeepSeek V4 Pro and Qwen3.8 Max on the Coding Agent Index while maintaining a lower cost than the latter. Overall, index-versus-cost efficiency metrics place Large 4 just outside the most attractive economic quadrant.
In speed evaluations, the model generated 116.1 tokens per second with a time-to-first-token of 1.46 seconds, running on Mistral's proprietary API infrastructure.
Infrastructure and Training
Training was conducted from scratch on 3,800 Nvidia Grace Blackwell GPUs housed in Mistral's European datacenters over approximately two months. This succeeds Large 3, which was trained on 3,000 Nvidia H200 chips and featured 675 billion total parameters.
Following a €3 billion Series D funding round, Mistral Compute plans to expand its infrastructure to 18,000 Grace Blackwell chips. The open-weight version of Large 4 is intended to run on private clouds or on-premises hardware once published.
Full architecture details and additional benchmarks are scheduled for release by the end of October. Until open weights arrive, benchmarks are treated as proprietary.
While European models continue to face competition from lower-cost alternatives, Mistral Large 4 demonstrates measurable progress in specialized enterprise applications.




