Veille Stratégique 09:00 EDT · Dimanche · Analyses multi-sources · 12 modèles croisés · Run #1 du jour
État au 30 août 2026 — Sources : ai.google.dev/gemini-api/docs/changelog + deepmind.google/models/gemini + deepmind.google/discover/blog + blog.google/innovation-and-ai
Données croisées : AA = Artificial Analysis Intel Index v4.1.1 (/63) · BL = BenchLM BenchAlign v5.2 (/100) · LB = LiveBench Overall (/100) · Coût/task AA depuis section "Cost per Intelligence Index Task". Prix in/out depuis OpenRouter (proxy multi-provider).
| # | Modèle | Compagnie | Date | Params | Params actifs | Prix in $/M | Prix out $/M | Compétences | Indice Intel /63 | Indice Code /100 | Coût/task AA ($) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (max) LEADER AA | Anthropic | 2026-07-24 | n/d | n/d | 10.00 | 50.00 | Ceiling capability, long-running agents, reasoning adaptatif max | 63 | 80.1 | n/d |
| 2 | Claude Opus 5 (xhigh) AA #1 | Anthropic | 2026-07-24 | n/d | n/d | 5.00 | 25.00 | Reasoning adaptatif, agentic, code, multimodal | 63 | 81.4 | 0.699 |
| 3 | Claude Fable 5 (max) LEADER LB | Anthropic | 2026-08-19 | n/d | n/d | 10.00 | 50.00 | Mythos-class, fallback Opus 4.8, long-running agents | 62 | 83.0 | 1.439 |
| 4 | Claude Opus 5 (high) | Anthropic | 2026-07-24 | n/d | n/d | 5.00 | 25.00 | Reasoning adaptatif, agentic, code, multimodal 1M ctx | 61 | 82.1 | 0.528 |
| 5 | GPT-5.6 Sol (max) | OpenAI | 2026-07-09 | n/d | n/d | 5.00 | 30.00 | Reasoning, agentic, code, multimodal 1.05M ctx | 61 | 81.0 | 0.515 |
| 6 | GPT-5.5 Thinking (xHigh) | OpenAI | 2026-06-XX | n/d | n/d | 5.00 | 30.00 | Thinking, reasoning xHigh effort, math 95.9 | 60 | 80.2 | 0.435 |
| 7 | Kimi K3 (max) OPEN #1 | Moonshot AI | 2026-07-16 | 2.8T | 104B | 3.00 | 15.00 | Reasoning, agentic, code, 1M context, open weights | 60 | 79.2 | 0.348 |
| 8 | Tencent Hy4 preview OPEN NEW 28/08 | Tencent Hunyuan | 2026-08-28 | 770B | 49B | 0.834 | 2.501 | MoE, coding agents, tool-use complexes, math/Blaschke-Lebesgue | n/d | n/d | n/d |
| 9 | Qwen 3.8 Max OPEN | Alibaba | 2026-08-12 | 2.4T | 95B (A95B) | 2.00 | 6.00 | Reasoning, agentic, code, 1.05M context, open weights MoE | 58 | 78.5 | 0.275 |
| 10 | GLM-5.3 (Z.AI) | Z.AI | 2026-08-18 | n/d | n/d | 1.40 | 4.40 | Reasoning, agentic, code, 1.05M context, reasoning always on | 58 | 76.1 | 0.450 |
| 11 | Grok 4.6 | xAI | 2026-08-06 | n/d | n/d | 2.00 | 6.00 | Reasoning, agentic, 500K ctx, multimodal audio | 57 | 78.0 | 0.207 |
| 12 | Gemini 3.7 Flash (High) | 2026-08-13 | n/d | n/d | 0.375 | 1.875 | Code, agents, multimodal, workhorse, 1M ctx, 50% off OpenRouter | 78.8 | 78.9 | 0.157 |
Divergences : Claude Opus 5 (max) AA: 63 saturé / BL: 82.94 / LB: 81.4 / Cost $0.699 · Claude Fable 5 (max) AA 62 / BL 82.96 #3 / LB Overall 83.0 #1 absolu / Cost $1.439 (plus cher) · Kimi K3 AA 60 #1 open weights / Cost $0.348 — équilibre prix/intel · Tencent Hy4 preview nouveau 28/08 — pas encore AA Intel publié (jeune) · DeepSeek V4 Pro 0813 cost $0.044/task = 16× moins cher que Opus 5 · Qwen3.8-27B AA 52 dense 27.78B, $0.094/task, single-GPU déployable · Gemini 3.7 Flash absent AA top mais BL 78.8 #1 value, 50% off OpenRouter jusqu'au 31/12/2026.
Référence OpenRouter : Seedream 5.0 Lite à $0.035/image (option budget), Seedream 5.0 Pro à partir de $0.045/image, Grok Imagine Image 2.0 sur xAI API de $0.02 (standard) à $0.08 (Quality Mode). Quality Mode rankings sur Text-to-Image Arena (Top 5 labos au 04/05/2026).