You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中同时实现系统音频的录音与实时语音转写?

解决系统音频实时录音+转写的多线程同步问题

原代码的核心问题是文件IO冲突和线程同步逻辑混乱,以下是针对性的解决方案:

核心问题分析

  • 录音和转写共用同一个音频文件,录音时会覆盖文件内容,转写过程中可能读取到不完整的音频数据,导致识别失败或程序崩溃。
  • 转写线程的sleep(2)和录音的5秒周期无关联,逻辑完全错位。
  • 未处理语音识别的异常场景(如识别失败、网络错误)。
  • 缺失time模块导入,代码无法正常运行。

修正方案

用内存队列替代文件作为音频数据的传递通道,让录音和转写线程完全解耦,同时添加异常处理保证稳定性:

import soundcard as sc
import soundfile as sf
import threading
import os
import speech_recognition as sr
import time
from queue import Queue

# 全局队列,用于传递录音数据
audio_queue = Queue(maxsize=5)  # 限制队列大小,避免内存溢出
CURRENT_DIR = os.path.dirname(os.path.abspath(__file__))
seconds = 5
speech = sr.Recognizer()

def record_buffer():
    samplerate = 44100
    # 初始化录音设备
    mic = sc.get_microphone(id=str(sc.default_speaker().name), include_loopback=True).recorder(samplerate=samplerate)
    with mic:
        while True:
            # 录制指定时长的音频
            audio_frames = mic.record(numframes=samplerate * seconds)
            # 将音频数据和采样率打包放入队列
            audio_queue.put((audio_frames, samplerate))
            # 短暂间隔避免队列满溢
            time.sleep(0.1)

def transcribe_audio():
    while True:
        if not audio_queue.empty():
            audio_frames, samplerate = audio_queue.get()
            # 临时保存音频到本地文件(用于speech_recognition读取)
            temp_file_path = os.path.join(CURRENT_DIR, "temp.wav")
            with sf.SoundFile(temp_file_path, mode='w', samplerate=samplerate, channels=2) as temp_file:
                temp_file.write(audio_frames)
            
            # 读取临时文件进行识别
            try:
                with sr.AudioFile(temp_file_path) as source:
                    audio_data = speech.record(source)
                    text = speech.recognize_google(audio_data, language="zh-CN")
                    print(f"识别结果: {text}")
            except sr.UnknownValueError:
                print("无法识别音频内容")
            except sr.RequestError as e:
                print(f"语音识别服务请求失败: {e}")
            finally:
                # 清理临时文件
                if os.path.exists(temp_file_path):
                    os.remove(temp_file_path)
            # 标记队列任务完成
            audio_queue.task_done()
        else:
            # 队列为空时休眠,降低CPU占用
            time.sleep(0.5)

# 启动录音线程(守护线程,随主线程退出)
record_thread = threading.Thread(target=record_buffer, daemon=True)
record_thread.start()

# 启动转写线程(守护线程)
transcribe_thread = threading.Thread(target=transcribe_audio, daemon=True)
transcribe_thread.start()

# 主线程保持运行,等待用户中断
try:
    while True:
        time.sleep(1)
except KeyboardInterrupt:
    print("程序已终止")

关键改进点

  • 使用Queue实现线程间安全的数据传递,彻底避免文件覆盖冲突。
  • 录音线程持续生产音频数据,转写线程异步消费,两者互不阻塞。
  • 添加完整的异常处理,覆盖识别失败、服务请求异常等场景。
  • 设置守护线程,保证主线程退出时子线程自动终止。
  • 临时文件用完即删,避免磁盘冗余数据。

内容的提问来源于stack exchange,提问作者CSR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:52:51