You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决MuseTalk实时唇形同步语音结束时的感知断层?

实时虚拟人唇形同步的过渡优化与模型选型问题

基础环境

  • 硬件:RTX 4090
  • 运行状态:25fps帧率实时唇形同步,基于aiohttp实现WebRTC推流
  • 技术依赖:基于Linly-Talker-Stream(LiveTalking分支)开发,核心模型为MuseTalk v1.5

管线工作机制

  • 语音阶段:MuseTalk根据输入音频生成新嘴部区域,融合到原始帧输出
  • 静音阶段:跳过MuseTalk处理,直接输出原始源帧

现存核心问题

语音结束瞬间,管线会在单帧内从「MuseTalk合成帧」切换到「原始源帧」。由于MuseTalk的已知局限——无法良好保留原始面部细节(如唇色、唇形),合成帧的嘴部区域明显比真实画面苍白、饱和度低,切换时会出现刺眼的颜色跳变。

已尝试的优化方案及结果

  • 放缓语音结束时的融合权重衰减速度(从0.05调整为0.02):跳变更明显,效果反向
  • 嘴部区域+周边窄带皮肤的LAB颜色匹配(通过肤色阈值过滤头巾/背景像素):脸颊区域颜色匹配改善,但嘴部仍偏白;增强校正力度会导致过饱和
  • 保持后溶解策略:缓存最后一帧说话画面,保持约240ms(6帧)后,用约160ms(4帧)交叉淡入到原始帧:弱化了边缘跳变,但颜色差异仍可被感知

核心需求

找到可隐藏语音结束过渡效果的方案,确保观众无法察觉虚拟人停止说话的瞬间,且必须满足25fps的直播实时性要求。

额外问题

  1. 是否存在能在语音阶段保留原始嘴部纹理细节的算法?避免MuseTalk合成画面的模糊问题
  2. 有没有其他可在RTX 4090上实时运行的高质量开源唇形同步模型?

融合函数代码

def get_image_blending(image, face, face_box, mask_array, crop_box, blend_strength=1.0):
    body = image
    x, y, x1, y1 = face_box
    x_s, y_s, x_e, y_e = crop_box
    face_large = copy.deepcopy(body[y_s:y_e, x_s:x_e])
    face_large[y-y_s:y1-y_s, x-x_s:x1-x_s] = face
    mask_image = cv2.cvtColor(mask_array, cv2.COLOR_BGR2GRAY)
    mask_image = (mask_image / 255).astype(np.float32) * blend_strength
    body[y_s:y_e, x_s:x_e] = cv2.blendLinear(
        face_large, body[y_s:y_e, x_s:x_e], mask_image, 1 - mask_image
    )
    return body

process_frames中的说话/静音分支(含保持后溶解尝试)

# --- SILENT branch ---
if audio_frames[0][1] != 0 and audio_frames[1][1] != 0:
    target_frame = self.frame_list_cycle[idx]

    if _was_speaking:
        # speech just ended — start hold phase using cached last composite
        if _last_speaking_composite is not None:
            _hold_counter = _hold_frames_total  # 6 frames (~240ms)
            _dissolve_counter = 0
    _was_speaking = False

    if _hold_counter > 0 and _last_speaking_composite is not None:
        combine_frame = _last_speaking_composite.copy()
        _hold_counter -= 1
        if _hold_counter == 0:
            _dissolve_counter = _dissolve_frames_total  # 4 frames (~160ms)
    elif _dissolve_counter > 0 and _last_speaking_composite is not None:
        alpha = 1.0 - (_dissolve_counter / float(_dissolve_frames_total))
        combine_frame = cv2.addWeighted(
            _last_speaking_composite, 1.0 - alpha, target_frame, alpha, 0
        )
        _dissolve_counter -= 1
    else:
        combine_frame = target_frame.copy()

# --- SPEAKING branch ---
else:
    _was_speaking = True
    current_frame = self.paste_back_frame(res_frame, idx, _blend_strength)
    combine_frame = current_frame
    # cache every speaking composite so we have it at speech-end
    _last_speaking_composite = combine_frame.copy()
    _hold_counter = 0
    _dissolve_counter = 0

问题截图说明

  • 最后一帧说话画面:嘴部区域苍白,饱和度低
  • 语音结束后的第一帧自然画面:色调偏暖,为真实面部色彩

内容的提问来源于stack exchange,提问作者Jimmy Fadel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 17:57:26