Google has introduced Live Avatar, a new capability for Gemini 3.8 Live that gives the company's conversational AI a near-real-time video face.在 WERSM 详细描述的官方声明中,谷歌将该功能展示为企业代理,可以听、看、说,同时保持动态的视觉角色——可通过该公司的商业人工智能产品 Gemini Enterprise 获得。
此次推出已经吸引了大约一周的报道:上周晚些时候,行业媒体开始报道 Gemini 3.8 Live 中的视觉呈现功能,《德干先驱报》周一进一步报道了这一消息,指出该功能“为人工智能聊天机器人增添了面孔”。对于评估 AI 工具和助手 的企业来说,该公告标志着企业对话式 AI 的发展方向:少一个聊天窗口,多一个永远在线的服务代表。
Live Avatar 实际上是做什么的
Live Avatar combines near-real-time video generation with Gemini's live dialogue capabilities.谷歌强调精确的口型同步、自然的表情和流畅的轮流——这些机制让说话的人感觉在场而不是神秘。
But the more consequential details sit behind the face.根据公告:
- 同时多模式输入。 代理可以同时处理视觉和音频输入,而不是分开处理它们。
- Asynchronous tool calls. Live Avatar can fetch information in the background while the conversation continues. Google's example is a hotel check-in, where the agent handles a complex task without leaving the guest staring at a loading state.
- 覆盖 97 种语言。 谷歌表示,该头像可以在 97 种语言之间移动,同时调整口型同步和表情,而不会出现明显的漂移——同一个品牌代理在不同市场上保持一致的身份,同时改变其口头表达方式。
That last point may be the feature's most commercially significant. A conventional chatbot makes the user wait for an answer; a Live Avatar is designed to keep the interaction moving while the system works underneath. The face creates continuity, but the underlying behaviour is closer to a service workflow that happens to remain conversational.
从聊天机器人到品牌形象
Google's framing asks brands to do more than enable a toggle.部署 Live Avatar 的公司被邀请设计一个角色——一种视觉识别、一种说话方式、一个客户将通过接触点和语言识别的角色。
这将企业人工智能从用户可以容忍的实用程序转变为与用户相关的存在,并将设计决策(代理的外观、反应方式、何时微笑)置于与底层模型质量相同的基础上。它还增加了一致性的风险:一个在不同语言中自相矛盾的化身,或者在不同渠道上表现不同的化身,都会破坏面孔本应建立的信任。
目前企业优先
据多份报告称,最初的可用性仅限于 Gemini Enterprise,谷歌尚未宣布向消费者推出的时间表。这使得 Live Avatar 处于 Google 人工智能功能的熟悉模式中:首先发布商业产品,在功能到达公共助理之前,部署受到控制并且货币化是直接的。
The enterprise-first approach also reflects what the feature demands.实时对话之上的实时视频生成的计算成本很高,而且企业工作负载(客户服务台、预订流程、面对面信息亭)具有明确的范围,可以根据其取代或增加的员工时间来证明成本的合理性。
信任方程
A face changes what users notice and what they forgive.谷歌强调精确的口型同步、自然的表情和流畅的轮流并非表面功夫:这些恰恰是人造角色断裂的接缝,用户发现延迟反应或不合时宜的表达比发现生硬的句子要快得多。
这对企业来说是双向的。 A well-executed avatar can make an automated service feel attended; a poorly executed one can make the same automation feel deceptive — a face performing attention that the underlying system does not have.从这个意义上说,谷歌描述的异步工具调用也是一种信任机制:在系统工作时,化身使对话保持明显的活跃状态,而不是假装答案已经准备好。
Enterprises deploying Live Avatar will effectively be judged on choreography — how the persona handles pauses, corrections and interruptions.该公告表明,谷歌已将这些时刻视为核心工程问题,而不是任何面向客户的部署所需要的润色。
为什么这很重要
语音助手让人工智能成为对话式的;头像使其更具表现力。如果 Live Avatar 的表现如所描述的那样——同步视觉和声音、背景工具使用、多语言口型同步——企业人工智能的界面就不再是文本框,而是开始成为公司控制的角色。
The open question is adoption: whether businesses actually want their AI to have a face, or whether the feature remains a demo-floor showpiece.谷歌押注于前者。 With 97 languages and background tool use built in, the company has made it easy for global brands to find out.
---
保持人工智能领先地位获取最新的人工智能新闻、分析和突破——尽在一个地方。
阅读更多人工智能新闻 →