You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Ruby中异步检查Google Speech API输出流的转录结果?

解决Google Speech流式识别实时返回结果及Async阻塞问题

问题根源

你的代码里output_stream.any?是阻塞式调用——Gapic的流对象在调用any?时,会一直等待直到有数据返回或者流关闭,这直接导致第二个异步任务被卡住,进而影响整个Async调度器的执行,连发送音频的任务也被阻塞。另外,默认的流式识别配置没有开启临时结果(interim results),这也是你只能在流结束后拿到最终结果的核心原因。

修复步骤

  • 开启临时结果配置:让Google Speech实时返回中间转录结果,而非仅在流结束后返回最终结果
  • 正确处理输出流:放弃loop + any?的阻塞式检查,直接迭代输出流,利用其异步可枚举特性自动处理结果
  • 优化异步任务结构:确保音频推送和结果处理的任务并行执行,互不干扰

修改后的代码

require 'google/cloud/speech'
require 'gapic-common'
require 'async'

audio_file_path = './output.webm'
speech = Google::Cloud::Speech.speech

audio_content = File.binread audio_file_path
bytes_total   = audio_content.size
bytes_sent    = 0
chunk_size    = 32_000

input_stream = Gapic::StreamInput.new
output_stream = speech.streaming_recognize input_stream

# 开启临时结果,支持实时返回中间转录
config = {
  config: {
    encoding:                 :WEBM_OPUS,
    sample_rate_hertz:        48_000,
    language_code:            "pl-PL",
    enable_word_time_offsets: true,
    enable_interim_results:   true # 关键:添加这一行开启实时临时结果
  }
}
input_stream.push streaming_config: config

Async do |task|
  # 推送音频块的异步任务
  task.async do
    while bytes_sent < bytes_total
      puts "Sending audio chunk"
      chunk = audio_content[bytes_sent, chunk_size]
      input_stream.push audio_content: chunk
      puts "Sent audio chunk (#{bytes_sent + chunk.size}/#{bytes_total})"
      bytes_sent += chunk_size
      sleep 1 # 模拟麦克风实时流的间隔
    end

    puts "Finished sending audio, closing input stream"
    input_stream.close
  end

  # 处理识别结果的异步任务
  task.async do
    puts "Waiting for recognition results..."
    # 直接迭代output_stream,有结果时自动触发回调
    output_stream.each do |response|
      response.results.each do |result|
        # 区分临时结果和最终结果
        result_type = result.is_final ? "最终结果" : "临时结果"
        puts "#{result_type}: #{result.alternatives.first.transcript}"
        # 如果需要处理单词时间偏移,可取消下面注释
        # result.alternatives.first.words.each do |word|
        #   puts "Word: #{word.word}, start: #{word.start_time.seconds}, end: #{word.end_time.seconds}"
        # end
      end
    end
    puts "Recognition stream closed"
  end
end

关键说明

  • enable_interim_results: 这个参数是实现实时转录的核心,开启后Google Speech会在音频流传输过程中不断返回中间识别结果,而非等流结束才返回最终结果
  • output_stream.each: Gapic的输出流是异步迭代器,each方法会在后台等待结果,不会阻塞整个Async调度器,能和音频推送任务并行执行
  • 避免阻塞式检查: 不要用any?、first这类会阻塞等待的方法,直接用迭代器处理结果是最安全的异步方式

Rails WebSocket场景适配提示

把上述逻辑整合到WebSocket控制器时:

  • 客户端通过WebSocket发送音频块时,直接推送到input_stream
  • 当output_stream返回结果时,通过WebSocket将转录文本实时推送给客户端
  • 注意在WebSocket连接关闭时,及时关闭input_stream,避免资源泄漏

内容的提问来源于stack exchange,提问作者rebelyer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 17:22:33