Veille Stratégique 09:00 EDT · Analyses multi-sources · 12 modèles croisés · Run #1 du jour
État au 27 août 2026 — Sources : ai.google.dev/gemini-api/docs/changelog + deepmind.google/models/gemini + blog.google/innovation-and-ai/models-and-research
Données croisées : AA = Artificial Analysis Intel Index v4.1.1 (/63) · BL = BenchLM BenchAlign v5.2 (/100) · LB = LiveBench Overall (/100) · Coût/task AA depuis section "Cost per Intelligence Index Task". Prix in/out depuis OpenRouter (proxy multi-provider).
| # | Modèle | Compagnie | Date | Params | Params actifs | Prix in $/M | Prix out $/M | Compétences | Indice Intel /63 | Indice Code /100 | Coût/task AA ($) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Mythos 5 LEADER | Anthropic | 2026-06-XX | n/d | n/d | 10.00 | 50.00 | Ceiling capability, long-running agents, reasoning | 63 | 83.26 | n/d |
| 2 | Claude Opus 5 (max) | Anthropic | 2026-07-24 | n/d | n/d | 5.00 | 25.00 | Reasoning adaptatif, agentic, code, multimodal | 63 | 81.4 | 0.699 |
| 3 | Claude Fable 5 (max) | Anthropic | 2026-08-19 | n/d | n/d | 10.00 | 50.00 | Mythos-class, fallback Opus 4.8, long-running agents | 62 | 86.0 | 1.439 |
| 4 | GPT-5.6 Sol (max) | OpenAI | 2026-07-09 | n/d | n/d | 5.00 | 30.00 | Reasoning, agentic, code, multimodal 1.05M ctx | 61 | 83.9 | 0.515 |
| 5 | Kimi K3 (max) OPEN | Moonshot AI | 2026-07-16 | 2.8T | 104B | 3.00 | 15.00 | Reasoning, agentic, code, 1M context, open weights | 60 | 81.4 | 0.348 |
| 6 | Grok 4.6 | xAI | 2026-08-12 | ~1.5T | n/d | 2.00 | 6.00 | Real-time, 500K context, GitHub Copilot natif | 75.38 | 76.8 | 0.207 |
| 7 | Gemini 3.7 Flash High | 2026-08-13 | n/d | n/d | 0.375 | 1.875 | Code, agents, multimodal, workhorse, 1M ctx | 78.8 | 78.9 | 0.157 | |
| 8 | Qwen 3.8 Max OPEN | Alibaba | 2026-08-03 | 2.4T | 95B | 2.00 | 6.00 | Reasoning, multimodal, 1M context, open weights | 58 | 72.9 | 0.275 |
| 9 | Z.ai GLM 5.3 OPEN | Z.AI | 2026-08-14 | 743B | 40B | 1.40 | 4.40 | Cyber defense, Terminal-Bench, reasoning, 1M ctx | 76.1 | 79.0 | 0.450 |
| 10 | DeepSeek V4 Pro 0813 OPEN | DeepSeek | 2026-08-13 | ~1.6T | 49B | 0.27 | 1.10 | Reasoning max effort, code, ultra-cheap peak/off-peak | 53 | 77.2 | 0.044 |
| 11 | Qwen 3.8 27B OPEN | Alibaba | 2026-08-14 | 27.78B | 27.78B | 0.35 | 2.75 | Local dense, multimodal, Apache 2.0, single-GPU | 52 | 75.7 | 0.094 |
| 12 | Gemini 3.1 Pro Preview High | 2026-02-19 | n/d | n/d | 2.00 | 10.00 | Reasoning, code, multimodal Pro tier | 77.0 | 76.5 | 0.286 |
Divergences : Claude Mythos 5 BL 83.26 #1 absolu mais AA Intel non publié (accès limité). Claude Opus 5 (max) AA: 63 (saturé) / BL: 82.94 / LB: 81.4 / Coût $0.699. Kimi K3 AA: 60 #1 open weights / LB 79.2 / Cost $0.348 — équilibre prix/intel. DeepSeek V4 Pro 0813 : cost $0.044/task = 16× moins cher que Opus 5. Qwen3.8-27B : AA 52 (dense 27.78B) matche GPT-5.6 Luna, $0.094/task, single-GPU déployable. Gemini 3.7 Flash absent AA top mais BL 78.8 #1 value, 50% off OpenRouter jusqu'au 31/12/2026.
Référence OpenRouter : Seedream 5.0 Lite à $0.035/image (option budget), Seedream 5.0 Pro à partir de $0.045/image, Grok Imagine Image 2.0 sur xAI API de $0.02 (standard) à $0.08 (Quality Mode). Quality Mode rankings sur Text-to-Image Arena (Top 5 labos au 04/05/2026).