You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Azure AI Foundry中Phi-4-multimodal-instruct进行音频转文本时触发Invalid input错误

使用Azure AI Foundry中Phi-4-multimodal-instruct进行音频转文本时触发Invalid input错误

我明白你在尝试用Phi-4-multimodal-instruct处理音频转文本时遇到了这个让人困惑的Invalid input错误,咱们一步步来排查并解决这个问题。

错误根源分析

从你的代码和报错来看,主要问题出在音频内容的构造方式和可能的端点URL格式上:

  1. 你手动构造了音频内容的字典,但Azure AI Inference SDK要求使用官方提供的强类型类(比如AudioContentItem和InputAudio)来传递音频数据,手动字典的结构可能不符合API的严格校验要求。
  2. 你的ENDPOINT_URL包含了/models后缀,这可能不是正确的推理端点格式。

修正后的完整代码

下面是调整后的可运行代码,我会标注关键修改点:

import base64
import os
from azure.ai.inference import ChatCompletionsClient
from azure.ai.inference.models import SystemMessage, UserMessage, TextContentItem, AudioContentItem, InputAudio
from azure.core.credentials import AzureKeyCredential
from azure.identity import DefaultAzureCredential

# Configuration - 注意修正端点URL,去掉/models后缀
ENDPOINT_URL = "https://myaiservices.inference.ai.azure.com"  # 正确的推理端点格式
MODEL_NAME = "phi-4-multimodal-instruct"
AUDIO_PATH = "harvard.wav"  # Path to a sample audio file

def test_audio_prompt(client, audio_path):
    """Test an audio input with a prompt."""
    print("\n=== Testing Audio Prompt ===")
    try:
        # Open and encode the audio file to base64
        with open(audio_path, "rb") as audio_file:
            base64_audio = base64.b64encode(audio_file.read()).decode("utf-8")

        # 关键修改1:使用SDK提供的InputAudio和AudioContentItem类,而非手动字典
        input_audio = InputAudio(
            data=base64_audio,
            mime_type="audio/wav"
        )
        audio_content_item = AudioContentItem(audio=input_audio)

        # Make the chat completions request
        response = client.complete(
            messages=[
                SystemMessage(content="You are an AI assistant for translating and transcribing audio clips."),
                UserMessage(content=[
                    TextContentItem(text="Please transcribe this audio snippet and print the result."),
                    audio_content_item  # 使用SDK类对象而非自定义字典
                ])
            ],
            model=MODEL_NAME
        )

        # Print the response
        print("Audio Response:", response.choices[0].message.content)
    except FileNotFoundError:
        print(f"Audio file not found: {audio_path}")
    except Exception as e:
        # 关键修改2:打印更详细的错误信息,方便排查
        print(f"Audio Prompt Error: {str(e)}")
        # 如果有响应对象,打印更多细节(比如API返回的具体错误)
        if hasattr(e, 'response'):
            print(f"Response details: {e.response.json()}")

def main():
    # 初始化客户端 - 注意端点URL已修正
    client = ChatCompletionsClient(
        endpoint=ENDPOINT_URL,
        credential=DefaultAzureCredential(),
        credential_scopes=["https://cognitiveservices.azure.com/.default"]
    )
    test_audio_prompt(client, AUDIO_PATH)

if __name__ == "__main__":
    main()

额外注意事项

  1. 音频文件格式验证:确保你的WAV文件是Phi-4-multimodal支持的格式,推荐使用16kHz采样率、单声道、PCM编码的WAV文件,避免因音频编码不兼容导致的错误。
  2. 端点URL确认:登录Azure AI Foundry,找到你的模型部署,复制正确的推理端点,格式应为https://<your-resource-name>.inference.ai.azure.com,不要包含/models或其他后缀。
  3. 权限验证:确保你的DefaultAzureCredential有访问该模型部署的权限,或者可以改用AzureKeyCredential直接传入API密钥来测试:
    client = ChatCompletionsClient(
        endpoint=ENDPOINT_URL,
        credential=AzureKeyCredential("your-api-key-here")
    )
    

如果调整后仍然有问题,建议打印出完整的异常响应细节(代码中已添加相关逻辑),这样能更精准地定位API拒绝请求的具体原因。

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 07:24:51