You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Microsoft认知API实现语音转文字?及Node.js代码排障

Hey there! Let's tackle your two questions about Microsoft's Speech-to-Text API step by step.

如何使用Microsoft认知API实现语音转文字功能?

First, let's cover the basics you'll need to get started, then walk through two common implementation approaches for Node.js:

Prerequisites

  • An Azure account (the free tier works great for testing)
  • A Speech resource created in the Azure portal—this gives you an API key and region (like eastus or westeurope) for authentication
  • Audio files that meet the API's requirements: WAV format with PCM encoding, 16kHz sample rate, and single channel (mono). These specs are the most reliable for consistent results.

Option 1: Use the Microsoft Speech SDK for Node.js

This is the simplest and most robust method, as the SDK handles authentication, audio streaming, and error handling out of the box.

  1. Install the SDK package:

    npm install microsoft-cognitiveservices-speech-sdk
    
  2. Basic code example to transcribe a local WAV file:

    const sdk = require("microsoft-cognitiveservices-speech-sdk");
    const fs = require("fs");
    
    // Replace with your Speech resource key and region
    const speechConfig = sdk.SpeechConfig.fromSubscription("YOUR_API_KEY", "YOUR_REGION");
    speechConfig.speechRecognitionLanguage = "zh-CN"; // Set your target language
    
    // Point to your local WAV file
    const audioConfig = sdk.AudioConfig.fromWavFileInput(fs.readFileSync("./your-audio-file.wav"));
    
    // Initialize the recognizer
    const recognizer = new sdk.SpeechRecognizer(speechConfig, audioConfig);
    
    // Handle successful recognition
    recognizer.recognized = (s, e) => {
        if (e.result.reason === sdk.ResultReason.RecognizedSpeech) {
            console.log(`转文字结果: ${e.result.text}`);
        } else if (e.result.reason === sdk.ResultReason.NoMatch) {
            console.log(`未识别到有效语音: ${e.result.noMatchDetails}`);
        }
    };
    
    // Handle errors or cancellations
    recognizer.canceled = (s, e) => {
        console.log(`识别流程被取消: ${e.reason}`);
        if (e.reason === sdk.CancellationReason.Error) {
            console.log(`错误详情: ${e.errorDetails}`);
        }
        recognizer.stopContinuousRecognitionAsync();
    };
    
    // Start the recognition process
    recognizer.startContinuousRecognitionAsync();
    

Option 2: Call the REST API Directly

If you prefer raw HTTP requests without the SDK, you can send the audio file directly to the Speech-to-Text endpoint. Here's a quick example using axios:

const axios = require("axios");
const fs = require("fs");

const apiKey = "YOUR_API_KEY";
const region = "YOUR_REGION";
const audioBuffer = fs.readFileSync("./your-audio-file.wav");

axios.post(
    `https://${region}.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1?language=zh-CN`,
    audioBuffer,
    {
        headers: {
            "Ocp-Apim-Subscription-Key": apiKey,
            "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000",
        },
    }
)
.then(response => {
    console.log("转文字结果:", response.data.DisplayText);
})
.catch(error => {
    console.error("请求失败:", error.response?.data || error.message);
});
Node.js上传WAV文件语音转文字出错排查

Since you didn't share the specific error code or your code snippet, let's go through the most common culprits:

  • Authentication mistakes: Double-check that your API key and region match exactly what's listed in your Azure Speech resource. Typos here are the #1 cause of failures! Also, confirm the resource isn't disabled or expired.
  • Audio format mismatch: The API is strict about specs. Use ffmpeg to convert your WAV to the required format if needed:
    ffmpeg -i input.wav -acodec pcm_s16le -ar 16000 -ac 1 output.wav
    
  • File reading issues: Ensure your code points to the correct file path, and the Node.js process has read permissions for that file. Log the file size (e.g., console.log(fs.statSync("./file.wav").size)) to confirm it's loading properly.
  • Incorrect headers (REST API only): If using the REST endpoint, make sure the Content-Type header includes the exact codec and samplerate (e.g., audio/wav; codecs=audio/pcm; samplerate=16000). Missing or wrong headers will trigger errors.
  • Quota limits: Check your Azure portal to see if you've hit the free tier's request limit. If so, you may need to upgrade to a paid tier or wait for the quota to reset.
  • Network restrictions: If you're behind a firewall or proxy, ensure it allows traffic to Azure Speech service endpoints (e.g., *.speech.microsoft.com).

If you can share your specific error message and a snippet of your code, I can help narrow down the issue even further!

内容的提问来源于stack exchange,提问作者omega

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:18:31