You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Faster-Whisper在Python中进行语音转写无输出问题求助

Faster-Whisper在Python中进行语音转写无输出问题求助

我完全懂你现在的困扰——录音按钮点下去看起来一切正常,但转写文本就是死活不出来对吧?结合你的代码,我整理了几个最可能的问题点和对应的修复方案,你可以一步步排查:


1. 剩余音频未处理,导致短录音无输出

你的代码里设置了攒够1秒音频才会触发转写(len(audio_buffer) * CHUNK >= FRAME_RATE),如果录音时间不足1秒,或者录音结束后缓冲区里还剩一些音频片段,这些内容就会被直接忽略,自然看不到输出。

修复方案:

修改speech_recognition函数,在录音停止后强制处理缓冲区里剩下的所有音频:

def speech_recognition():
    """Processes audio from the queue and transcribes it using Faster-Whisper."""
    audio_buffer = []
    
    while is_recording or not recordings.empty():
        if not recordings.empty():
            data = recordings.get()
            audio_buffer.append(np.frombuffer(data, dtype=np.int16))
            # 处理满1秒的音频片段
            if len(audio_buffer) * CHUNK >= FRAME_RATE:
                audio_chunk = np.concatenate(audio_buffer).astype(np.float32) / 32768  
                audio_buffer = []
                segments, _ = model.transcribe(audio_chunk, language="en", beam_size=5)            
                with output:
                    for segment in segments:
                        display(segment.text)
    # 录音结束后,处理缓冲区剩余的所有音频
    if audio_buffer:
        audio_chunk = np.concatenate(audio_buffer).astype(np.float32) / 32768  
        segments, _ = model.transcribe(audio_chunk, language="en", beam_size=5)            
        with output:
            for segment in segments:
                display(segment.text)

2. PyAudio实例重复创建,引发设备资源冲突

你在全局和record_microphone函数里都创建了p = pyaudio.PyAudio()实例,这会导致音频设备被重复占用,甚至出现隐性的录制失败(看起来在录音,但实际没拿到音频数据)。

修复方案:

删除函数内的PyAudio实例创建,复用全局的实例,同时不要在函数内终止全局实例(否则后续无法再次录音):

def record_microphone():
    """Records audio from the microphone and puts it in a queue."""
    global is_recording
    
    # 去掉这里的p = pyaudio.PyAudio(),用全局的p
    stream = p.open(format=AUDIO_FORMAT, channels=CHANNELS, rate=FRAME_RATE,
                    input=True, input_device_index=default_device_index, frames_per_buffer=CHUNK)

    while is_recording:
        data = stream.read(CHUNK)
        recordings.put(data)
            
    stream.stop_stream()
    stream.close()
    # 移除这行:p.terminate()

如果需要清理资源,可以在Notebook的最后单独执行:

p.terminate()

3. 确认音频设备是否正确录制

有时候默认输入设备可能不是你正在使用的麦克风,导致录制的是静音数据,自然转写不出内容。

排查方案:

在start_recording函数里添加设备信息打印,确认当前使用的麦克风是否正确:

def start_recording(data):
    """Starts recording and transcription threads."""
    global is_recording
    is_recording = True

    with output:
        display("Listening...")
        # 打印当前使用的输入设备信息
        device_info = p.get_device_info_by_index(default_device_index)
        display(f"当前使用麦克风:{device_info['name']}")
        display(f"录制参数:{CHANNELS}声道,{FRAME_RATE}Hz采样率")
    
    record_thread = Thread(target=record_microphone)
    transcribe_thread = Thread(target=speech_recognition)

    record_thread.start()
    transcribe_thread.start()

4. Jupyter线程输出刷新问题

Jupyter Notebook对线程内的display支持有时候不太稳定,可能转写完成了但输出没刷新出来。

优化方案:

改用print或者结合clear_output强制刷新输出:

from IPython.display import clear_output

# 在转写输出的地方修改
with output:
    clear_output(wait=True)  # 先清空旧输出
    for segment in segments:
        print(segment.text)  # 用print替代display,更稳定

你可以先从第1和第2点开始排查,这两个是最常见的原因,应该能解决大部分问题。

备注:内容来源于stack exchange,提问作者Nik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 09:00:27