You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

替换ElevenLabs为Edge-TTS遇UnboundLocalError,求问题排查与修复

文本转MP3功能迁移至Edge-TTS的错误修复

问题背景

原本基于ElevenLabs API实现的文本转MP3功能运行正常,但原型阶段成本过高,遂改用Edge-TTS,运行时出现以下错误:

UnboundLocalError: cannot access local variable 'audio_segment' where it is not associated with a value

原ElevenLabs API代码

from elevenlabs import generate, set_api_key, Voice, VoiceSettings
from pydub import AudioSegment
import io
import os
import hashlib
from utils.os_stuff import get_env_var_or_fail
import logging

set_api_key(get_env_var_or_fail('ELEVEN_LABS_API_KEY'))

HOST_VOICE = Voice(
    voice_id="21m00Tcm4TlvDq8ikWAM",
    name="Rachel",
    category="premade",
    settings=VoiceSettings(stability=0.35, similarity_boost=0.9),
)

ADS_VOICE = Voice(
    voice_id="TxGEqnHWrfWFTfGW9XjX",
    name="Josh",
    category="premade",
    settings=VoiceSettings(stability=0.35, similarity_boost=0.9),
)


def load_audio_bytes(audio_bytes):
    audio_file = io.BytesIO(audio_bytes)
    audio_segment = AudioSegment.from_file(audio_file, format='mp3')
    return audio_segment


def convert_text_to_mp3(text, voice):
    # Generate the cache key by MD5 hashing the text
    cache_key = hashlib.md5(text.encode()).hexdigest()

    # Check if the file already exists in cache
    cache_dir = ".eleven_labs_cache"
    cache_file = os.path.join(cache_dir, f"{cache_key}.mp3")

    if not os.path.exists(cache_dir):
        os.makedirs(cache_dir)

    if os.path.exists(cache_file):
        # If it does exist, load and return as an AudioSegment
        audio_segment = AudioSegment.from_mp3(cache_file)
    else:
        # If it does not exist, call the API, create, save, and return as an AudioSegment
        char_count = len(text)
        logging.info(f'Calling eleven labs for {char_count} chars...')
        section_1_voice_over = load_audio_bytes(generate(
            text=text,
            voice=voice
        ))
        section_1_voice_over.export(cache_file, format='mp3')
        audio_segment = section_1_voice_over

    return audio_segment

错误的Edge-TTS代码

import asyncio
import edge_tts
from pydub import AudioSegment

VOICE = "en-GB-SoniaNeural"

def convert_text_to_mp3(text):
    loop = asyncio.get_event_loop_policy().get_event_loop()
    try:
        audio_segment = loop.run_until_complete(edge_tts.Communicate(text, VOICE))
    finally:
        loop.close()
        return audio_segment

错误原因分析

  1. edge_tts.Communicate返回值错误处理:该方法返回的是异步生成器,而非直接的音频数据,直接赋值给audio_segment无法得到有效音频内容。
  2. 异常场景下变量未初始化:如果run_until_complete抛出异常,audio_segment未被赋值,但finally块直接返回该变量,触发UnboundLocalError。
  3. 缺失缓存逻辑:原代码通过本地缓存避免重复API调用,新代码未实现该功能,会增加不必要的请求和成本。

修复方案

以下是修复后的完整代码,保留原缓存逻辑,正确处理Edge-TTS异步音频流,并修复变量初始化问题:

import asyncio
import edge_tts
from pydub import AudioSegment
import io
import os
import hashlib
import logging

# 配置Edge-TTS语音映射,对应原ElevenLabs的语音
VOICES = {
    "Rachel": "en-GB-SoniaNeural",
    "Josh": "en-US-JasonNeural"
}

async def fetch_edge_tts_audio(text, voice):
    """异步获取Edge-TTS的完整音频字节流"""
    communicate = edge_tts.Communicate(text, voice)
    audio_bytes = b""
    async for chunk in communicate.stream():
        if chunk["type"] == "audio":
            audio_bytes += chunk["data"]
    return audio_bytes

def convert_text_to_mp3(text, voice_name="Rachel"):
    # 生成缓存Key,区分不同语音避免冲突
    cache_key = hashlib.md5(f"{text}_{voice_name}".encode()).hexdigest()
    cache_dir = ".edge_tts_cache"
    cache_file = os.path.join(cache_dir, f"{cache_key}.mp3")

    # 确保缓存目录存在
    if not os.path.exists(cache_dir):
        os.makedirs(cache_dir)

    audio_segment = None
    try:
        if os.path.exists(cache_file):
            # 读取缓存文件
            audio_segment = AudioSegment.from_mp3(cache_file)
            logging.info(f"Loaded cached audio from {cache_file}")
        else:
            # 调用Edge-TTS获取音频
            char_count = len(text)
            logging.info(f'Calling Edge-TTS for {char_count} chars with voice {voice_name}...')
            voice = VOICES.get(voice_name, VOICES["Rachel"])
            audio_bytes = asyncio.run(fetch_edge_tts_audio(text, voice))
            
            # 转换为AudioSegment并保存到缓存
            audio_file = io.BytesIO(audio_bytes)
            audio_segment = AudioSegment.from_file(audio_file, format='mp3')
            audio_segment.export(cache_file, format='mp3')
            logging.info(f"Saved audio to cache: {cache_file}")
        
        return audio_segment
    except Exception as e:
        logging.error(f"Error converting text to MP3: {str(e)}")
        raise

关键修复点说明

  • 正确处理异步音频流:通过async for迭代communicate.stream()获取音频字节块,拼接成完整的音频数据。
  • 初始化变量避免未绑定错误:提前定义audio_segment = None,确保无论是否出现异常,变量都有合法定义。
  • 保留并优化缓存逻辑:沿用MD5缓存策略,同时加入语音名称区分缓存,避免不同语音的相同文本缓存冲突。
  • 兼容原函数调用方式:保留voice_name参数,与原ElevenLabs函数的调用逻辑兼容,减少代码改动量。
  • 增加异常处理:加入日志记录和异常抛出,便于调试和问题定位。

内容的提问来源于stack exchange,提问作者rajtilakjee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 10:53:29