Unofficial AI-summarized news site (not affiliated with any AI company)
AI News JP / www.ai-news.jp
🟠 Important AI Summary · Source: DeepSeek Release Notes

KV cache drops to 1/8 size — direct hit on agent execution costs. V4-Pro phasing out.

DeepSeek Launches V4.1-Flash: 8x Faster Inference, Lower Costs, Phasing Out V4-Pro

Original: DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient | DeepSeek API Docs

Importance: 本番 API の強制ルーティング変更とコスト構造大幅改善で、既存 DeepSeek ユーザーに即座に影響

Summary

DeepSeek unveiled V4.1-Flash, a 552B-parameter MoE model featuring a new Causal Encoder–Decoder architecture with only 8B active parameters for input and 16B for output. The model reduces KV cache storage requirements to 1/8 of the previous generation, significantly cutting agent execution costs while delivering benchmark performance exceeding V4-Pro across speed, cost, and capability. Starting September 14, 2026, all V4-Pro API requests will route to V4.1-Flash at V4.1-Flash pricing until V4.1-Pro launches. Official partners WorkBuddy and OpenCode now fully support the new model.

Key Points

  • KV cache shrinks to 1/8, slashing agent execution costs
  • V4-Pro sunset; auto-routing to V4.1-Flash begins Sept 14
  • Outperforms V4-Pro on benchmarks across speed/cost/capability
  • 8B/16B active params enable faster inference
  • WorkBuddy, OpenCode, other partners already supporting it
View developer notes (APIs, breaking changes, migration)

V4.1-Flash technical specs: 552B-parameter MoE with new Causal Encoder–Decoder enabling 8B active params at input, 16B at output. KV cache compressed to 1/8 vs prior gen (reduces SSD storage burden). Set model to `deepseek-flash`. Legacy `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` temporarily route to V4.1-Flash. Effective Sept 14 2026 04:00 UTC, all `deepseek-v4-pro` requests auto-route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro GA. Cache-hit charges often dominate agent costs; compression directly reduces per-inference overhead.

モデルAPI/SDKパフォーマンスAudience: 開発者Audience: 企業導入担当

Source: https://api-docs.deepseek.com/news/news260910

Outlet: DeepSeek Release Notes

This article is an AI-generated summary (OpenAI GPT-4o-mini) of publicly available information from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Sakana, and other vendors. The original source URL is always provided in accordance with fair-use citation requirements. Summaries are AI-generated and may contain mistranslations or misinterpretations. Always verify details with the original source.