You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中WAV录音与语音转文本异常求助:录制文件无内容

Fixing Empty Recording & Speech-to-Text Failure in Python

Hey there! Let's troubleshoot why your output.wav is coming up empty and get that speech recognition working smoothly. Here are the most common fixes and a revised version of your code:

Common Causes & Solutions

1. Microphone Permissions

First off, make sure your Python environment has access to your microphone—this is a super common oversight:

  • Windows: Head to Settings > Privacy & Security > Microphone, and check that the terminal/code editor running your script has permission.
  • macOS: Open System Settings > Privacy & Security > Microphone, and tick the box next to your terminal or IDE.
  • Linux: Add your user to the audio group with sudo usermod -aG audio $USER, then log out and back in to apply the change.

2. Incorrect Channel Configuration

Most built-in microphones are mono (1 channel), but your code uses channels=2 (stereo). If your mic only outputs one channel, recording in stereo can result in silent files. Switch to mono:

myrecording = sd.rec(int(seconds * fs), samplerate=fs, channels=1)

3. Pick the Right Microphone Device

If you have multiple audio devices (like a webcam mic + headset), sounddevice might be picking the wrong one. List all available devices first to find your microphone's index:

print(sd.query_devices())

Then specify the device in your recording line (replace 0 with your mic's index):

myrecording = sd.rec(int(seconds * fs), samplerate=fs, channels=1, device=0)

4. Verify Recording Before Saving

Add a quick playback step right after recording to confirm audio is actually being captured:

sd.wait()
print("Playing back your recording...")
sd.play(myrecording, fs)
sd.wait()

5. Tweak Speech Recognition Settings

Sometimes ambient noise calibration is too aggressive, or the language isn't specified properly. Adjust these parts of your code:

with sr.AudioFile(sound) as source:
    # Shorten ambient noise adjustment (default is 1 second) to avoid cutting your speech
    recognizer.adjust_for_ambient_noise(source, duration=0.5)
    print("Converting the answer to text...")
    audio = recognizer.listen(source)

try:
    # Specify your language if it's not English (e.g., 'en-US', 'fr-FR')
    text = recognizer.recognize_google(audio, language='en-US')
    print("The converted text:", text)
except sr.UnknownValueError:
    print("Oops! Couldn't understand the audio—try speaking more clearly.")
except sr.RequestError as e:
    print(f"Failed to connect to Google's service: {e}")
except Exception as e:
    print(f'An unexpected error occurred: {e}')

Revised Full Code

import speech_recognition as sr
import sounddevice as sd
import numpy as np
from scipy.io.wavfile import write

fs = 44100  # Sample rate
seconds = 15  # Duration of recording

# List available devices to find your microphone index
print("Available audio devices:")
print(sd.query_devices())
device_index = int(input("Enter your microphone device index: "))

print("Start recording the answer.....")
# Record with mono channel and specified device
myrecording = sd.rec(int(seconds * fs), samplerate=fs, channels=1, device=device_index)
sd.wait()  # Wait until recording is finished

# Play back to verify recording
print("Playing back your recording...")
sd.play(myrecording, fs)
sd.wait()

# Save as WAV file
write('output.wav', fs, myrecording.astype(np.int16))

recognizer = sr.Recognizer()
sound = "output.wav"

with sr.AudioFile(sound) as source:
    recognizer.adjust_for_ambient_noise(source, duration=0.5)
    print("Converting the answer to text...")
    audio = recognizer.listen(source)

try:
    text = recognizer.recognize_google(audio, language='en-US')
    print("The converted text:", text)
except sr.UnknownValueError:
    print("Oops! Couldn't understand the audio. Make sure you're speaking clearly into the mic.")
except sr.RequestError as e:
    print(f"Failed to connect to Google Speech Recognition: {e}")
except Exception as e:
    print(f'An error occurred: {e}')

Quick Extra Tips

  • Speak close to the microphone and minimize background noise while recording.
  • If issues persist, try updating your libraries: pip install --upgrade speechrecognition sounddevice scipy numpy
  • Test with a different microphone (like a headset) to rule out hardware issues.

内容的提问来源于stack exchange,提问作者Hirushi Ekanayake

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 21:42:53