使用Azure AI Foundry中Phi-4-multimodal-instruct进行音频转文本时触发Invalid input错误
使用Azure AI Foundry中Phi-4-multimodal-instruct进行音频转文本时触发Invalid input错误
我明白你在尝试用Phi-4-multimodal-instruct处理音频转文本时遇到了这个让人困惑的Invalid input错误,咱们一步步来排查并解决这个问题。
错误根源分析
从你的代码和报错来看,主要问题出在音频内容的构造方式和可能的端点URL格式上:
- 你手动构造了音频内容的字典,但Azure AI Inference SDK要求使用官方提供的强类型类(比如
AudioContentItem和InputAudio)来传递音频数据,手动字典的结构可能不符合API的严格校验要求。 - 你的ENDPOINT_URL包含了
/models后缀,这可能不是正确的推理端点格式。
修正后的完整代码
下面是调整后的可运行代码,我会标注关键修改点:
import base64 import os from azure.ai.inference import ChatCompletionsClient from azure.ai.inference.models import SystemMessage, UserMessage, TextContentItem, AudioContentItem, InputAudio from azure.core.credentials import AzureKeyCredential from azure.identity import DefaultAzureCredential # Configuration - 注意修正端点URL,去掉/models后缀 ENDPOINT_URL = "https://myaiservices.inference.ai.azure.com" # 正确的推理端点格式 MODEL_NAME = "phi-4-multimodal-instruct" AUDIO_PATH = "harvard.wav" # Path to a sample audio file def test_audio_prompt(client, audio_path): """Test an audio input with a prompt.""" print("\n=== Testing Audio Prompt ===") try: # Open and encode the audio file to base64 with open(audio_path, "rb") as audio_file: base64_audio = base64.b64encode(audio_file.read()).decode("utf-8") # 关键修改1:使用SDK提供的InputAudio和AudioContentItem类,而非手动字典 input_audio = InputAudio( data=base64_audio, mime_type="audio/wav" ) audio_content_item = AudioContentItem(audio=input_audio) # Make the chat completions request response = client.complete( messages=[ SystemMessage(content="You are an AI assistant for translating and transcribing audio clips."), UserMessage(content=[ TextContentItem(text="Please transcribe this audio snippet and print the result."), audio_content_item # 使用SDK类对象而非自定义字典 ]) ], model=MODEL_NAME ) # Print the response print("Audio Response:", response.choices[0].message.content) except FileNotFoundError: print(f"Audio file not found: {audio_path}") except Exception as e: # 关键修改2:打印更详细的错误信息,方便排查 print(f"Audio Prompt Error: {str(e)}") # 如果有响应对象,打印更多细节(比如API返回的具体错误) if hasattr(e, 'response'): print(f"Response details: {e.response.json()}") def main(): # 初始化客户端 - 注意端点URL已修正 client = ChatCompletionsClient( endpoint=ENDPOINT_URL, credential=DefaultAzureCredential(), credential_scopes=["https://cognitiveservices.azure.com/.default"] ) test_audio_prompt(client, AUDIO_PATH) if __name__ == "__main__": main()
额外注意事项
- 音频文件格式验证:确保你的WAV文件是Phi-4-multimodal支持的格式,推荐使用16kHz采样率、单声道、PCM编码的WAV文件,避免因音频编码不兼容导致的错误。
- 端点URL确认:登录Azure AI Foundry,找到你的模型部署,复制正确的推理端点,格式应为
https://<your-resource-name>.inference.ai.azure.com,不要包含/models或其他后缀。 - 权限验证:确保你的DefaultAzureCredential有访问该模型部署的权限,或者可以改用AzureKeyCredential直接传入API密钥来测试:
client = ChatCompletionsClient( endpoint=ENDPOINT_URL, credential=AzureKeyCredential("your-api-key-here") )
如果调整后仍然有问题,建议打印出完整的异常响应细节(代码中已添加相关逻辑),这样能更精准地定位API拒绝请求的具体原因。
内容来源于stack exchange
相关产品推荐
相关产品推荐

