You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Discord机器人MP3转文本功能:无法接收附件的解决方案咨询

解决方案

问题分析

你的代码存在两个核心问题:

  • 斜杠命令模式下无法获取附件:ctx.message在纯斜杠命令调用时不存在,导致无法读取附件;
  • OpenAI API调用错误:你误用了文本补全的Completion接口,而Whisper语音转文本需要专用的Audio.transcribe接口。

修复步骤与完整代码

1. 兼容混合命令的附件接收

为hybrid命令添加attachment参数(支持斜杠命令直接上传),同时保留引用消息获取附件的逻辑,覆盖两种使用场景:

  • 直接在命令后上传MP3(斜杠/前缀命令都支持)
  • 引用包含MP3的消息来转录

2. 修正OpenAI Whisper API调用

使用openai.Audio.transcribe接口,指定Whisper模型(如whisper-1),并正确传入音频文件对象。

以下是修改后的完整代码:

import io
import openai
from discord.ext import commands
from discord import Attachment, Message
import checks  # 假设你的checks模块已正确导入

class Transcribe(commands.Cog, name="transcribe"):
    def __init__(self, bot):
        self.bot = bot
        # 确保已设置OpenAI API密钥
        # openai.api_key = "你的API密钥"

    @commands.hybrid_command(
        name="transcribe",
        description="将MP3音频文件转录为文本",
    )
    @checks.not_blacklisted()
    @checks.is_owner()
    async def transcribe(self, ctx, attachment: Attachment = None, message: Message = None):
        # 优先处理直接传入的附件
        target_attachment = attachment
        
        # 如果没有直接传附件,检查引用的消息
        if not target_attachment:
            # 检查是否引用了消息
            if message:
                if message.attachments:
                    target_attachment = message.attachments[0]
                else:
                    await ctx.send("引用的消息中没有附件")
                    return
            # 再检查前缀命令的消息附件
            elif ctx.message and ctx.message.attachments:
                target_attachment = ctx.message.attachments[0]
            else:
                await ctx.send("请直接上传MP3文件,或引用包含MP3的消息")
                return

        # 验证文件格式是否为MP3
        if not target_attachment.filename.lower().endswith(".mp3"):
            await ctx.send("仅支持MP3格式的音频文件")
            return

        # 下载音频内容
        try:
            audio_content = await target_attachment.read()
        except Exception as e:
            await ctx.send(f"下载文件失败: {str(e)}")
            return

        # 使用BytesIO包装音频内容,供Whisper API读取
        audio_file = io.BytesIO(audio_content)
        audio_file.name = target_attachment.filename

        # 调用OpenAI Whisper API
        try:
            response = openai.Audio.transcribe(
                model="whisper-1",
                file=audio_file,
                language="zh"  # 可根据需要指定语言,比如"en"
            )
            transcription = response["text"].strip()
        except Exception as e:
            await ctx.send(f"转录失败: {str(e)}")
            return

        # 返回转录结果
        if transcription:
            await ctx.send(f"转录结果:\n{transcription}")
        else:
            await ctx.send("未能识别音频内容")

# 记得将Cog添加到机器人中
# bot.add_cog(Transcribe(bot))

关键说明

  • 附件兼容:通过添加attachment: Attachment参数,斜杠命令会自动生成上传文件的选项;同时支持传入message: Message参数来引用带附件的消息,覆盖更多使用场景。
  • Whisper API调用:使用whisper-1模型,这是OpenAI专为语音转文本优化的模型,支持多种语言。BytesIO用于将下载的音频内容转换成API可读取的文件对象。
  • 错误处理:添加了文件格式验证、下载失败、转录失败的异常捕获,提升稳定性。

内容的提问来源于stack exchange,提问作者Ammad Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 05:47:42