Discord机器人MP3转文本功能:无法接收附件的解决方案咨询
解决方案
问题分析
你的代码存在两个核心问题:
- 斜杠命令模式下无法获取附件:
ctx.message在纯斜杠命令调用时不存在,导致无法读取附件; - OpenAI API调用错误:你误用了文本补全的
Completion接口,而Whisper语音转文本需要专用的Audio.transcribe接口。
修复步骤与完整代码
1. 兼容混合命令的附件接收
为hybrid命令添加attachment参数(支持斜杠命令直接上传),同时保留引用消息获取附件的逻辑,覆盖两种使用场景:
- 直接在命令后上传MP3(斜杠/前缀命令都支持)
- 引用包含MP3的消息来转录
2. 修正OpenAI Whisper API调用
使用openai.Audio.transcribe接口,指定Whisper模型(如whisper-1),并正确传入音频文件对象。
以下是修改后的完整代码:
import io import openai from discord.ext import commands from discord import Attachment, Message import checks # 假设你的checks模块已正确导入 class Transcribe(commands.Cog, name="transcribe"): def __init__(self, bot): self.bot = bot # 确保已设置OpenAI API密钥 # openai.api_key = "你的API密钥" @commands.hybrid_command( name="transcribe", description="将MP3音频文件转录为文本", ) @checks.not_blacklisted() @checks.is_owner() async def transcribe(self, ctx, attachment: Attachment = None, message: Message = None): # 优先处理直接传入的附件 target_attachment = attachment # 如果没有直接传附件,检查引用的消息 if not target_attachment: # 检查是否引用了消息 if message: if message.attachments: target_attachment = message.attachments[0] else: await ctx.send("引用的消息中没有附件") return # 再检查前缀命令的消息附件 elif ctx.message and ctx.message.attachments: target_attachment = ctx.message.attachments[0] else: await ctx.send("请直接上传MP3文件,或引用包含MP3的消息") return # 验证文件格式是否为MP3 if not target_attachment.filename.lower().endswith(".mp3"): await ctx.send("仅支持MP3格式的音频文件") return # 下载音频内容 try: audio_content = await target_attachment.read() except Exception as e: await ctx.send(f"下载文件失败: {str(e)}") return # 使用BytesIO包装音频内容,供Whisper API读取 audio_file = io.BytesIO(audio_content) audio_file.name = target_attachment.filename # 调用OpenAI Whisper API try: response = openai.Audio.transcribe( model="whisper-1", file=audio_file, language="zh" # 可根据需要指定语言,比如"en" ) transcription = response["text"].strip() except Exception as e: await ctx.send(f"转录失败: {str(e)}") return # 返回转录结果 if transcription: await ctx.send(f"转录结果:\n{transcription}") else: await ctx.send("未能识别音频内容") # 记得将Cog添加到机器人中 # bot.add_cog(Transcribe(bot))
关键说明
- 附件兼容:通过添加
attachment: Attachment参数,斜杠命令会自动生成上传文件的选项;同时支持传入message: Message参数来引用带附件的消息,覆盖更多使用场景。 - Whisper API调用:使用
whisper-1模型,这是OpenAI专为语音转文本优化的模型,支持多种语言。BytesIO用于将下载的音频内容转换成API可读取的文件对象。 - 错误处理:添加了文件格式验证、下载失败、转录失败的异常捕获,提升稳定性。
内容的提问来源于stack exchange,提问作者Ammad Ali
相关产品推荐
相关产品推荐

