You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Discord语音双向翻译机器人依赖问题及代码优化求助

Discord语音双向翻译机器人依赖问题及代码优化求助

我正在开发一个能在Discord语音频道实现翻译功能的机器人,用Whisper做语音转文字、NLLB做文本翻译、gTTS做文字转语音。现在遇到了卡壳的依赖问题,同时也想请大家帮忙看看代码在性能、稳定性、功能扩展性上有没有可以优化的地方。

先说最头疼的依赖问题:我已经安装了discord.py[voice],但代码里导入discord.sinks的时候一直提示找不到这个模块,试了重新安装、升级discord.py都没用,也找不到单独的discord.sinks库可以安装,这该怎么解决啊?

下面是我目前写的代码,麻烦大家帮忙看看:

import os
import asyncio
import discord
from discord.ext import commands
from discord.sinks import MP3Sink  # 这里导入一直报错
from transformers import (
    WhisperForConditionalGeneration,
    WhisperProcessor,
    AutoTokenizer,
    AutoModelForSeq2SeqLM
)
from gtts import gTTS
from io import BytesIO
import logging

# 配置日志
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# 配置信息(请替换为自己的参数)
DISCORD_TOKEN = "---------------------"
VOICE_CHANNEL_ID = 937370496989806607  # 目标语音频道ID
TEXT_CHANNEL_ID = 937370496989806605  # 日志/翻译文本输出频道ID

# 模型配置
WHISPER_MODEL_NAME = "openai/whisper-large-v2"
NLLB_MODEL_NAME = "facebook/nllb-200-distilled-600M"

# 固定语言设置(暂时写死,想改成可动态切换的)
SOURCE_LANG = "eng_Latn"
TARGET_LANG = "rus_Cyrl"

class VoiceTranslateBot(commands.Bot):
    def __init__(self, **kwargs):
        intents = discord.Intents.all()
        super().__init__(command_prefix='!', intents=intents)
        self.voice_client = None
        self.whisper_processor = None
        self.whisper_model = None
        self.nllb_tokenizer = None
        self.nllb_model = None
        self.audio_queue = asyncio.Queue()

    async def setup_hook(self):
        await self.load_models()
        await self.connect_to_voice()
        self.loop.create_task(self.audio_processing_loop())

    async def load_models(self):
        logger.info("正在加载AI模型...")
        self.whisper_processor = WhisperProcessor.from_pretrained(WHISPER_MODEL_NAME)
        self.whisper_model = WhisperForConditionalGeneration.from_pretrained(WHISPER_MODEL_NAME).to('cuda')
        self.nllb_tokenizer = AutoTokenizer.from_pretrained(NLLB_MODEL_NAME)
        self.nllb_model = AutoModelForSeq2SeqLM.from_pretrained(NLLB_MODEL_NAME).to('cuda')
        logger.info("所有模型加载完成。")

    async def connect_to_voice(self):
        channel = self.get_channel(VOICE_CHANNEL_ID)
        if channel:
            self.voice_client = await channel.connect()
            # 启动音频录制
            self.voice_client.start_recording(
                MP3Sink(), self.on_voice_data, self.loop
            )
            logger.info("已连接到目标语音频道并开始录制。")
        else:
            logger.warning("找不到指定ID的语音频道!")

    async def on_voice_data(self, sink, audio_data, *args):
        """收到语音片段时的回调函数"""
        for user_id, data in audio_data.items():
            if user_id != self.user.id:  # 忽略机器人自己的声音
                await self.audio_queue.put(data.file.read())

    async def audio_processing_loop(self):
        """异步处理队列中的音频数据"""
        while True:
            audio_bytes = await self.audio_queue.get()
            try:
                text = await self.transcribe_audio(audio_bytes)
                if text.strip():
                    await self.process_translation(text)
            except Exception as e:
                logger.error(f"音频处理出错:{str(e)}")

    async def transcribe_audio(self, audio_bytes):
        """用Whisper将语音转成文本"""
        # 现在是先写临时文件再处理,感觉有点浪费IO,能不能直接用字节流?
        with open("temp_input.wav", "wb") as f:
            f.write(audio_bytes)
        from transformers import pipeline
        transcriber = pipeline(
            "automatic-speech-recognition",
            model=self.whisper_model,
            tokenizer=self.whisper_processor,
            device=0
        )
        result = transcriber("temp_input.wav")
        os.remove("temp_input.wav")
        return result['text']

    def translate_text(self, text):
        """用NLLB将文本翻译成目标语言"""
        self.nllb_tokenizer.src_lang = SOURCE_LANG
        encoded = self.nllb_tokenizer(text, return_tensors="pt").to('cuda')
        forced_bos_token_id = self.nllb_tokenizer.lang_code_to_id[TARGET_LANG]
        generated_ids = self.nllb_model.generate(**encoded, forced_bos_token_id=forced_bos_token_id)
        return self.nllb_tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]

    def synthesize_speech(self, text, lang='ru'):
        """用gTTS将翻译后的文本转成语音"""
        tts = gTTS(text=text, lang=lang)
        fp = BytesIO()
        tts.write_to_fp(fp)
        fp.seek(0)
        return fp

    async def process_translation(self, text):
        """完整的翻译流程:转录→翻译→播报"""
        translated = self.translate_text(text)
        logger.info(f"翻译结果:{translated}")
        # 可以在这里把原文和译文发到文本频道
        text_channel = self.get_channel(TEXT_CHANNEL_ID)
        if text_channel:
            await text_channel.send(f"原文:{text}\n译文:{translated}")
        audio_fp = self.synthesize_speech(translated)
        await self.play_audio(audio_fp)

    async def play_audio(self, audio_fp):
        """在语音频道播放翻译后的语音"""
        source = discord.FFmpegPCMAudio(audio_fp, pipe=True)
        if self.voice_client.is_playing():
            self.voice_client.stop()
        self.voice_client.play(source)

    async def on_ready(self):
        logger.info(f"机器人已启动,当前登录身份:{self.user}")

# 启动机器人
bot = VoiceTranslateBot()
bot.run(DISCORD_TOKEN)

具体求助内容

1. 紧急依赖问题

执行pip install discord.py[voice]后,from discord.sinks import MP3Sink始终报错ModuleNotFoundError: No module named 'discord.sinks'。试过升级discord.py到最新版、重新创建虚拟环境安装,都没用,也找不到单独的discord.sinks包。请问该怎么解决这个导入问题?

2. 代码优化建议

希望大家能从以下几个方面帮忙提提改进意见:

  • 模型资源管理:现在启动时直接把模型全加载到GPU,有没有更高效的方式?比如延迟加载、模型量化减少显存占用?
  • 音频IO优化:当前转录时要写临时文件,能不能直接用字节流喂给Whisper,避免磁盘IO开销?
  • 队列稳定性:用异步队列存音频数据,会不会出现队列积压、音频丢失的情况?有没有更好的队列管理策略?
  • 错误处理完善:当前错误处理比较简单,有没有需要覆盖的边界情况(比如语音频道断开、模型推理失败、音频格式错误、用户频繁说话导致的处理不过来)?
  • 功能扩展性:现在是固定源语言和目标语言,怎么改成可以用命令动态切换语言、甚至自动识别说话人的语言?
  • 性能提升:如果多个用户同时说话,当前单线程处理会不会卡顿?有没有办法实现并行处理?

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.07 10:20:28