Gemini gets human-like expression. Text-to-avatar shift could be a turning point for AI UX
Introducing Gemini 3.8 Live with Live Avatar
Original: Introducing Gemini 3.8 Live with Live Avatar
Importance: マルチモーダルAIの大型機能追加で、ユーザー体験と開発機能セットが大きく拡張される
Summary
Google DeepMind announced Live Avatar support for Gemini 3.8 Live, enabling AI to interact via dynamic avatar representations in real time. This evolution from text and static interfaces aims to deliver more natural, human-like interactions for education, customer support, and entertainment use cases, marking a significant step in multimodal AI advancement.
Key Points
- Live Avatar now available in Gemini 3.8 Live
- Dynamic avatars enhance conversational experience vs. text-only
- Targets education, support, entertainment verticals
- Real-time streaming at API level
View developer notes (APIs, breaking changes, migration)
Live Avatar in Gemini 3.8 Live likely extends the Gemini API with new streaming endpoints for avatar rendering, introducing avatar_mode parameters and real-time facial expression synthesis. Developers will control via REST/WebSocket, integrating camera/mic input. Compatibility with existing text and video modes expected; latency targets <200ms RTT. SDK rollout anticipated for Python, Node.js, Go; full API docs pending.
Source: https://deepmind.google/blog/introducing-gemini-38-live-with-live-avatar/
Outlet: Google DeepMind
This article is an AI-generated summary (OpenAI GPT-4o-mini) of publicly available information from Anthropic, OpenAI, Google, Meta, Mistral, DeepSeek, Sakana, and other vendors. The original source URL is always provided in accordance with fair-use citation requirements. Summaries are AI-generated and may contain mistranslations or misinterpretations. Always verify details with the original source.