You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab中麦克风音频采集至SpeechRecognition库识别的问题:设备不可用及音频格式报错

Alright, let's tackle your Google Colab audio recognition issues head-on—you've got two key problems here: accessing your local microphone in the cloud environment, and feeding the audio directly to SpeechRecognition without saving a file first. Let's break this down step by step.

1. Fixing the "No Default Input Device Available" Error

Google Colab runs on remote servers, so it can't directly access your local microphone. We need to use your browser's media API to capture audio, then send that data to the Colab runtime. Here's a script that adds a recording button to your notebook and captures audio as base64-encoded data:

from IPython.display import display, Javascript
from google.colab import output
import base64

def record_audio():
  display(Javascript('''
    const recordBtn = document.createElement('button');
    recordBtn.textContent = 'Start Recording';
    document.body.appendChild(recordBtn);

    const audioChunks = [];
    const mediaRecorder = new MediaRecorder(navigator.mediaDevices.getUserMedia({ audio: true }));

    mediaRecorder.ondataavailable = event => {
      audioChunks.push(event.data);
    };

    mediaRecorder.onstop = () => {
      const audioBlob = new Blob(audioChunks, { type: 'audio/wav' });
      const reader = new FileReader();
      reader.onload = () => {
        const base64Audio = reader.result.split(',')[1];
        google.colab.kernel.invokeFunction('notebook.handle_audio', [base64Audio], {});
      };
      reader.readAsDataURL(audioBlob);
    };

    recordBtn.onclick = () => {
      if (mediaRecorder.state === 'recording') {
        mediaRecorder.stop();
        recordBtn.textContent = 'Start Recording';
      } else {
        mediaRecorder.start();
        recordBtn.textContent = 'Stop Recording';
      }
    };
  '''))

# Store captured audio data
captured_audio = []

def handle_audio_data(data):
  captured_audio.append(data)

output.register_callback('notebook.handle_audio', handle_audio_data)

Run this code, then call the record_audio() function—you'll see a button to start/stop recording. When you stop, the audio data will be stored in the captured_audio list.

2. Directly Feeding Audio to SpeechRecognition (No File Save)

SpeechRecognition works best with 16-bit PCM WAV, 16kHz sample rate, mono audio. The browser's captured audio might not match these specs, so we'll use pydub to convert it in memory, then create an AudioData object directly (no file write required).

First, install pydub if you haven't already:

!pip install pydub

Now process the captured audio and run recognition:

import speech_recognition as sr
import io
from pydub import AudioSegment

# Decode the base64 audio to bytes
audio_bytes = base64.b64decode(captured_audio[0])

# Convert audio to SpeechRecognition-compatible format in memory
audio_segment = AudioSegment.from_file(io.BytesIO(audio_bytes), format='wav')
# Set to 16-bit sample width, 16kHz rate, mono channel
audio_segment = audio_segment.set_sample_width(2).set_frame_rate(16000).set_channels(1)

# Export to an in-memory BytesIO object
audio_io = io.BytesIO()
audio_segment.export(audio_io, format='wav')
audio_io.seek(0)

# Run recognition
r = sr.Recognizer()
with sr.AudioFile(audio_io) as source:
    audio_data = r.record(source)

try:
    recognized_text = r.recognize_google(audio_data)
    print(f"Recognized Text: {recognized_text}")
except sr.UnknownValueError:
    print("Google Speech Recognition couldn't understand the audio")
except sr.RequestError as e:
    print(f"Error connecting to Google's service: {e}")

Why Your Earlier WAV File Failed

The error you saw happened because the browser's captured WAV likely used a non-compatible format (e.g., 32-bit float instead of 16-bit PCM, or a different sample rate). SpeechRecognition is strict about input specs, so converting with pydub ensures the audio matches what the library expects.

内容的提问来源于stack exchange,提问作者Hasanen A. Sahib

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:27:29