Veille Stratégique 09:00 EDT · Lundi · Analyses multi-sources · 12 modèles croisés · Run #1 du jour
État au 31 août 2026 — Sources : ai.google.dev/gemini-api/docs/changelog + deepmind.google/models/gemini + deepmind.google/discover/blog + blog.google/innovation-and-ai
Données croisées : AA = Artificial Analysis Intel Index v4.1.1 (/63) · BL = BenchLM BenchAlign v5.2 (/100) · LB = LiveBench Overall (/100) · Coût/task AA depuis section "Cost per Intelligence Index Task". Prix in/out depuis OpenRouter (proxy multi-provider).
| # | Modèle | Compagnie | Date | Params | Params actifs | Prix in $/M | Prix out $/M | Compétences | Indice Intel /63 | Indice Code /100 | Coût/task AA ($) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max) LEADER AA | Anthropic | 2026-07-24 | n/d | n/d | 10.00 | 50.00 | Ceiling capability, long-running agents, reasoning adaptatif max, 1M ctx | 63 | 80.1 | 2.34 |
| 2 | Claude Fable 5 (max) LEADER LB | Anthropic | 2026-06-09 | n/d | n/d | 10.00 | 50.00 | Mythos-class, fallback Opus 4.8, biology/cyber safeguards, 1M ctx | 62 | 83.0 | 3.14 |
| 3 | GPT-5.6 Sol (max) | OpenAI | 2026-07-09 | n/d | n/d | 5.00 | 30.00 | Reasoning, agentic, code, multimodal 1.05M ctx, default Plus/Pro ChatGPT | 61 | 81.0 | 1.23 |
| 4 | Grok 4.6 (high) | xAI / SpaceXAI | 2026-08-12 | n/d | n/d | 2.00 | 6.00 | Reasoning, agentic, 500K ctx, multimodal, $0.84/task AA cheapest frontier | 61 | 78.0 | 0.84 |
| 5 | Kimi K3 (max) OPEN #1 | Moonshot AI | 2026-07-16 | 2.8T | 104B | 3.00 | 15.00 | Reasoning, agentic, code, 1M ctx, open weights, Frontend Code Arena #1 | 60 | 79.2 | 0.84 |
| 6 | GLM-5.3 (max) | Z.AI / Zhipu | 2026-08-14 | ~750B | ~40B | 1.40 | 4.40 | Reasoning always on, coding, defensive security, 1M ctx, post-training only | 60 | 76.1 | 0.68 |
| 7 | Qwen 3.8 Max OPEN | Alibaba | 2026-08-03 | 2.4T | 95B (A95B) | 2.00 | 6.00 | Reasoning, agentic, code, 1M context, multimodal, open weights MoE | 58 | 78.5 | 1.13 |
| 8 | Qwen3.8 2.4T A95B OPEN | Alibaba | 2026-08-12 | 2.4T | 95B | 2.00 | 6.00 | Open weights text-only checkpoint of Qwen3.8-Max, 262K ctx native | 58 | 75.3 | 0.95 |
| 9 | Gemini 3.7 Flash (high) | 2026-08-13 | n/d | n/d | 0.75 | 3.75 | Workhorse intelligent, code, agents, multimodal, 1M ctx, ~340 t/s | 56 | 78.8 | 0.40 | |
| 10 | DeepSeek V4 Pro 0813 OPEN | DeepSeek | 2026-08-13 | n/d | n/d | 0.66 | 1.98 | Coding, agentic, 1M ctx, MIT, peak/off-peak pricing, ultra-cheap | 55 | 77.4 | 0.044 |
| 11 | Seed 2.1 Turbo | ByteDance Seed | 2026-08-12 | n/d | n/d | 0.50 | 2.50 | Reasoning, agentic, 262K ctx, multimodal, performance/prix | n/d | 77.0 | n/d |
| 12 | Muse Spark 1.2 (xhigh) | Meta | 2026-08-05 | n/d | n/d | 1.25 | 4.25 | Multimodal, agentic, 1.05M ctx, open-weight strategy Meta | 57 | 78.0 | 0.40 |
Divergences : Claude Opus 5 (max) AA: 63 saturé / LB: 80.1 / Cost $2.34 · Claude Fable 5 (max) AA 62 / LB Overall 83.0 #1 absolu / Cost $3.14 (plus cher) · Kimi K3 AA 60 #1 open weights / Cost $0.84 · GLM-5.3 AA 60, Cost $0.68 — le moins cher de la frontier cluster · DeepSeek V4 Pro 0813 cost $0.044/task = 53× moins cher qu'Opus 5 · Gemini 3.7 Flash AA 56 (top workhorse) · Grok 4.6 AA 61 / Cost $0.84 tied #2 value frontier derrière GLM-5.3.
Référence OpenRouter : Seedream 5.0 Lite à $0.035/image (option budget), Seedream 5.0 Pro à partir de $0.045/image, Grok Imagine Image 2.0 sur xAI API de $0.02 (standard) à $0.08 (Quality Mode). Quality Mode rankings sur Text-to-Image Arena (Top 5 labos au 04/05/2026).