Unofficial AI-summarized news site (not affiliated with any AI company)
AI News JP / www.ai-news.jp
🔵 Standard AI Summary · Source: DeepSeek Release Notes

Multimodal agent performance hits Opus 4.8 levels at V4-Flash pricing — text + vision combo now viable

DeepSeek Launches V4-Flash-Vision-Exp: Multimodal API Now Available

Original: DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now Live | DeepSeek API Docs

Importance: 新しいマルチモーダルモデルの一般向けリリースで、APIユーザーに対する選択肢拡大と機能強化をもたらすが、既存運用への即座の破壊的影響はない

Summary

DeepSeek releases DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model now live on its API platform. It matches V4-Flash on text tasks but makes major strides on multimodal agent benchmarks, approaching Opus 4.8 performance. Images are tokenized at up to 384 tokens each using V4-Flash pricing. Supports mixed text+image input via base64, external URLs, or Files API; images can be uploaded once and reused via file_id. Works seamlessly with agent frameworks for broader practical workflows.

Key Points

  • DeepSeek-V4-Flash-Vision-Exp now GA on API
  • Multimodal agent performance approaches Opus 4.8
  • Images tokenized (up to 384 tokens each)
  • Files API enables image reuse, bandwidth savings
  • Compatible with major agent frameworks
View developer notes (APIs, breaking changes, migration)

Model identifier: deepseek-v4-flash-vision-exp. Supports Chat Completions, Messages & Responses APIs. Images tokenized at up to 384 tokens each at V4-Flash pricing. Files API integration: upload once, reference by file_id across requests (saves bandwidth, enables reuse). Image input via base64, external URLs, or Files API. DeepSeek Harness 0.1.1 released with native support. Set model='deepseek-v4-flash-vision-exp'. Visual understanding combined with tool integrations unlocks broader agent framework workflows.

API/SDKモデルパフォーマンスAudience: 開発者Audience: 企業導入担当

Source: https://api-docs.deepseek.com/news/news260821

Outlet: DeepSeek Release Notes

This article is an AI-generated summary (OpenAI GPT-4o-mini) of publicly available information from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Sakana, and other vendors. The original source URL is always provided in accordance with fair-use citation requirements. Summaries are AI-generated and may contain mistranslations or misinterpretations. Always verify details with the original source.