You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP Laravel调用Azure发音评估API返回AccuracyScore为0的问题排查

Azure发音评估API返回AccuracyScore始终为0.0的问题排查

在PHP Laravel应用中对接Azure Pronunciation Assessment服务时,API返回200成功响应,但返回结果的AccuracyScore始终为0.0。而在Azure Speech Studio中测试相同发音,准确率可达90%以上。怀疑问题出在音频文件处理环节,以下是实现细节及排查方向:

前端Vue音频录制代码

methods: {
    recordAudio() {
      navigator.mediaDevices.getUserMedia({audio: true, video: false})
        .then(stream => {
          this.mediaRecorder = new MediaRecorder(stream);
          this.mediaRecorder.addEventListener('start', this.onRecordingStart);
          this.mediaRecorder.addEventListener('stop', this.onRecordingStop);
          this.mediaRecorder.addEventListener('dataavailable', this.onRecordingDataAvailable);
          this.mediaRecorder.start();
        })
        .catch(error => {
          console.log(error);
        });
    },
    stopRecording() {
      this.mediaRecorder.stop();
    },
    onRecordingStart() {
      this.isRecording = true;
    },
    onRecordingDataAvailable(event) {
      this.audioChunks.push(event.data);
    },
    onRecordingStop() {
      this.isRecording = false;
      const audioBlob = new Blob(this.audioChunks, {'type': 'audio/wav'});
      this.assessPronunciation(audioBlob);
      const audioUrl = URL.createObjectURL(audioBlob);
      this.audio = new Audio(audioUrl);
    },
    assessPronunciation(audioBlob) {
      const formData = new FormData();
      formData.append('audio', audioBlob, 'recording.wav');
      formData.append('text', this.text);
      axios.post('/api/pronunciation-assessment', formData)
        .then(res => {
        })
        .catch(err => {
          console.log(err);
        });
    },
}

后端Laravel控制器代码

public function apiPostPronunciationAssessment(
        Request $request,
        AzureSpeechServicesApiClient $speechClient
    ): string {
        $audio = $request->file('audio');
        $text = $request->get('text');

        return $speechClient->assessPronunciation($text, $audio->getContent());
    }

Azure API客户端代码

<?php

namespace App\Services\Speech;

use GuzzleHttp\Client;
use GuzzleHttp\RequestOptions;
use Illuminate\Config\Repository;

class AzureSpeechServicesApiClient
{
    private string $key;
    private string $region;
    private string $pronunciationEndpoint;
    private Client $client;

    public function __construct(Repository $config)
    {
        $this->key = $config->get('services.azureSpeech.key');
        $this->region = $config->get('services.azureSpeech.location');
        $this->pronunciationEndpoint =
            "https://$this->region.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1?language=:lang";
    }

    public function assessPronunciation(string $text, string $audio): string
    {
        $response = $this->client()->post(
            $this->pronunciationEndpoint(),
            [
                RequestOptions::HEADERS => $this->pronunciationHeaders($text),
                RequestOptions::BODY => $audio
            ]
        );

        return $response->getBody()->getContents();
    }

    public function region(): string
    {
        return $this->region;
    }

    private function client(): Client
    {
        if (!isset($this->client)) {
            $this->client = new Client();
        }

        return $this->client;
    }

    private function pronunciationHeaders(string $text): array
    {
        return [
            'Ocp-Apim-Subscription-Key' => $this->key,
            'Content-Type' => 'audio/wav',
            'Accept' => 'application/json;text/xml',
            'Pronunciation-Assessment' => base64_encode(json_encode([
                'ReferenceText' => $text,
                'GradingSystem' => 'HundredMark',
                'PhonemeAlphabet' => 'IPA',
            ])),

        ];
    }

    private function pronunciationEndpoint(): string
    {
        $language = targetLang() === "en" ? "en-US" : "es-ES";

        return str_replace(':lang', $language, $this->pronunciationEndpoint);
    }
}

API返回结果示例

{
  "RecognitionStatus": "Success",
  "Offset": 5700000,
  "Duration": 1100000,
  "NBest": [
    {
      "Confidence": 0.84944737,
      "Lexical": "crook",
      "ITN": "crook",
      "MaskedITN": "crook",
      "Display": "Crook.",
      "AccuracyScore": 0.0,
      "Words": [
        {
          "Word": "crook",
          "Offset": 5700000,
          "Duration": 1100000,
          "Confidence": 0.0,
          "AccuracyScore": 0.0,
          "Syllables": [...],
          "Phonemes": [...]
        }
      ]
    }
  ],
  "DisplayText": "Crook."
}

排查方向

  • 音频格式不符合要求:Azure发音评估要求WAV音频为16kHz采样率、16位单声道、PCM编码。MediaRecorder默认生成的Blob可能不满足该条件,可通过ffmpeg -i recording.wav检查本地保存的音频参数。若不符,前端需调整录制配置(如指定audio约束)或使用Recorder.js等工具生成标准WAV。
  • ReferenceText匹配问题:确保Pronunciation-Assessment头中的ReferenceText与实际发音文本完全一致(大小写、标点、空格均需匹配),语言参数需和文本语言对应。
  • 音频传输损坏:对比前端生成的Blob文件与后端保存的文件,检查是否在传输或读取过程中出现二进制数据损坏。可在后端将接收的音频保存为文件,验证是否可正常播放且参数正确。
  • 端点与头参数验证:确认发音评估端点URL正确,解码Pronunciation-Assessment头的base64值,检查JSON结构是否符合Azure要求(如GradingSystem、PhonemeAlphabet参数是否合法)。

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 02:14:53