Discord语音双向翻译机器人依赖问题及代码优化求助
Discord语音双向翻译机器人依赖问题及代码优化求助
我正在开发一个能在Discord语音频道实现翻译功能的机器人,用Whisper做语音转文字、NLLB做文本翻译、gTTS做文字转语音。现在遇到了卡壳的依赖问题,同时也想请大家帮忙看看代码在性能、稳定性、功能扩展性上有没有可以优化的地方。
先说最头疼的依赖问题:我已经安装了discord.py[voice],但代码里导入discord.sinks的时候一直提示找不到这个模块,试了重新安装、升级discord.py都没用,也找不到单独的discord.sinks库可以安装,这该怎么解决啊?
下面是我目前写的代码,麻烦大家帮忙看看:
import os import asyncio import discord from discord.ext import commands from discord.sinks import MP3Sink # 这里导入一直报错 from transformers import ( WhisperForConditionalGeneration, WhisperProcessor, AutoTokenizer, AutoModelForSeq2SeqLM ) from gtts import gTTS from io import BytesIO import logging # 配置日志 logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) # 配置信息(请替换为自己的参数) DISCORD_TOKEN = "---------------------" VOICE_CHANNEL_ID = 937370496989806607 # 目标语音频道ID TEXT_CHANNEL_ID = 937370496989806605 # 日志/翻译文本输出频道ID # 模型配置 WHISPER_MODEL_NAME = "openai/whisper-large-v2" NLLB_MODEL_NAME = "facebook/nllb-200-distilled-600M" # 固定语言设置(暂时写死,想改成可动态切换的) SOURCE_LANG = "eng_Latn" TARGET_LANG = "rus_Cyrl" class VoiceTranslateBot(commands.Bot): def __init__(self, **kwargs): intents = discord.Intents.all() super().__init__(command_prefix='!', intents=intents) self.voice_client = None self.whisper_processor = None self.whisper_model = None self.nllb_tokenizer = None self.nllb_model = None self.audio_queue = asyncio.Queue() async def setup_hook(self): await self.load_models() await self.connect_to_voice() self.loop.create_task(self.audio_processing_loop()) async def load_models(self): logger.info("正在加载AI模型...") self.whisper_processor = WhisperProcessor.from_pretrained(WHISPER_MODEL_NAME) self.whisper_model = WhisperForConditionalGeneration.from_pretrained(WHISPER_MODEL_NAME).to('cuda') self.nllb_tokenizer = AutoTokenizer.from_pretrained(NLLB_MODEL_NAME) self.nllb_model = AutoModelForSeq2SeqLM.from_pretrained(NLLB_MODEL_NAME).to('cuda') logger.info("所有模型加载完成。") async def connect_to_voice(self): channel = self.get_channel(VOICE_CHANNEL_ID) if channel: self.voice_client = await channel.connect() # 启动音频录制 self.voice_client.start_recording( MP3Sink(), self.on_voice_data, self.loop ) logger.info("已连接到目标语音频道并开始录制。") else: logger.warning("找不到指定ID的语音频道!") async def on_voice_data(self, sink, audio_data, *args): """收到语音片段时的回调函数""" for user_id, data in audio_data.items(): if user_id != self.user.id: # 忽略机器人自己的声音 await self.audio_queue.put(data.file.read()) async def audio_processing_loop(self): """异步处理队列中的音频数据""" while True: audio_bytes = await self.audio_queue.get() try: text = await self.transcribe_audio(audio_bytes) if text.strip(): await self.process_translation(text) except Exception as e: logger.error(f"音频处理出错:{str(e)}") async def transcribe_audio(self, audio_bytes): """用Whisper将语音转成文本""" # 现在是先写临时文件再处理,感觉有点浪费IO,能不能直接用字节流? with open("temp_input.wav", "wb") as f: f.write(audio_bytes) from transformers import pipeline transcriber = pipeline( "automatic-speech-recognition", model=self.whisper_model, tokenizer=self.whisper_processor, device=0 ) result = transcriber("temp_input.wav") os.remove("temp_input.wav") return result['text'] def translate_text(self, text): """用NLLB将文本翻译成目标语言""" self.nllb_tokenizer.src_lang = SOURCE_LANG encoded = self.nllb_tokenizer(text, return_tensors="pt").to('cuda') forced_bos_token_id = self.nllb_tokenizer.lang_code_to_id[TARGET_LANG] generated_ids = self.nllb_model.generate(**encoded, forced_bos_token_id=forced_bos_token_id) return self.nllb_tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0] def synthesize_speech(self, text, lang='ru'): """用gTTS将翻译后的文本转成语音""" tts = gTTS(text=text, lang=lang) fp = BytesIO() tts.write_to_fp(fp) fp.seek(0) return fp async def process_translation(self, text): """完整的翻译流程:转录→翻译→播报""" translated = self.translate_text(text) logger.info(f"翻译结果:{translated}") # 可以在这里把原文和译文发到文本频道 text_channel = self.get_channel(TEXT_CHANNEL_ID) if text_channel: await text_channel.send(f"原文:{text}\n译文:{translated}") audio_fp = self.synthesize_speech(translated) await self.play_audio(audio_fp) async def play_audio(self, audio_fp): """在语音频道播放翻译后的语音""" source = discord.FFmpegPCMAudio(audio_fp, pipe=True) if self.voice_client.is_playing(): self.voice_client.stop() self.voice_client.play(source) async def on_ready(self): logger.info(f"机器人已启动,当前登录身份:{self.user}") # 启动机器人 bot = VoiceTranslateBot() bot.run(DISCORD_TOKEN)
具体求助内容
1. 紧急依赖问题
执行pip install discord.py[voice]后,from discord.sinks import MP3Sink始终报错ModuleNotFoundError: No module named 'discord.sinks'。试过升级discord.py到最新版、重新创建虚拟环境安装,都没用,也找不到单独的discord.sinks包。请问该怎么解决这个导入问题?
2. 代码优化建议
希望大家能从以下几个方面帮忙提提改进意见:
- 模型资源管理:现在启动时直接把模型全加载到GPU,有没有更高效的方式?比如延迟加载、模型量化减少显存占用?
- 音频IO优化:当前转录时要写临时文件,能不能直接用字节流喂给Whisper,避免磁盘IO开销?
- 队列稳定性:用异步队列存音频数据,会不会出现队列积压、音频丢失的情况?有没有更好的队列管理策略?
- 错误处理完善:当前错误处理比较简单,有没有需要覆盖的边界情况(比如语音频道断开、模型推理失败、音频格式错误、用户频繁说话导致的处理不过来)?
- 功能扩展性:现在是固定源语言和目标语言,怎么改成可以用命令动态切换语言、甚至自动识别说话人的语言?
- 性能提升:如果多个用户同时说话,当前单线程处理会不会卡顿?有没有办法实现并行处理?
内容来源于stack exchange
相关产品推荐
相关产品推荐

