Laravel 11中如何整合Google Cloud Speech-to-Text与Pusher并解决超时问题?
问题分析与解决方案
当前方案的核心问题
- HTTP无状态特性不适合实时音频流:你用
fetch发送的每一块音频都是独立的HTTP请求,后端控制器每次实例化都会新建一个Google Speech流式连接。Google Speech的流式API需要连续、实时的音频输入,而零散的HTTP请求会导致每个流式连接只收到一块音频后就断流,触发超时错误。 - 控制器构造函数的资源复用错误:控制器的构造函数在每次HTTP请求时都会执行,意味着每个音频块都对应一个全新的Google Speech流,完全无法形成连续的音频会话。
优化架构
改用WebSocket长连接维持前端与后端的实时通信:
- 前端通过WebSocket持续发送音频块,保持连接打开
- 后端为每个WebSocket连接创建一个独立的Google Speech流式会话,持续接收音频并处理
- 转录结果通过Pusher实时广播回前端,满足实时翻译需求
具体实现代码
前端代码修改(WebSocket替代HTTP)
let mediaRecorder; let socket; document.getElementById('record').addEventListener('click', async () => { const stream = await navigator.mediaDevices.getUserMedia({ audio: true }); // 明确指定编码格式,与后端配置对齐 mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/webm; codecs=opus' }); // 建立WebSocket连接(生产环境用wss,开发环境用ws) socket = new WebSocket(`${window.location.protocol === 'https:' ? 'wss' : 'ws'}://${window.location.host}/ws/speech-stream`); socket.onopen = () => { console.log('音频流连接已建立'); mediaRecorder.start(250); // 每250ms发送一块音频 }; mediaRecorder.ondataavailable = event => { // 确保连接处于打开状态再发送数据 if (socket.readyState === WebSocket.OPEN) { socket.send(event.data); } }; mediaRecorder.onstop = () => { socket.close(); console.log('音频流连接已关闭'); }; }); document.getElementById('stop').addEventListener('click', () => { mediaRecorder.stop(); }); // Pusher接收转录结果逻辑保留,优化展示格式 Pusher.logToConsole = true; var pusher = new Pusher('{{ env('PUSHER_APP_KEY') }}', { cluster: 'eu' }); var channel = pusher.subscribe('speech-transcript-created-channel'); var pusherTextArea = document.getElementById('pusherTextArea'); channel.bind('speech-transcript-created', function(data) { pusherTextArea.value += data.transcript + '\n'; });
后端WebSocket控制器(处理音频流与Google Speech交互)
<?php namespace App\Http\Controllers\WebSocket; use App\Events\SpeechTranscriptCreated; use Google\Cloud\Speech\V1\RecognitionConfig; use Google\Cloud\Speech\V1\StreamingRecognitionConfig; use Google\Cloud\Speech\V1\StreamingRecognizeRequest; use Google\Cloud\Speech\V1\SpeechClient; use Ratchet\ConnectionInterface; use Ratchet\MessageComponentInterface; class SpeechStreamController implements MessageComponentInterface { // 存储每个连接对应的Google Speech资源 protected $speechClients = []; protected $streams = []; public function onOpen(ConnectionInterface $conn) { // 初始化Google Speech客户端 $speechClient = new SpeechClient([ 'transport' => 'grpc', 'credentials' => base_path(env('GOOGLE_APPLICATION_CREDENTIALS')) ]); // 配置流式识别,启用临时结果(实时返回) $config = new StreamingRecognitionConfig([ 'config' => new RecognitionConfig([ 'encoding' => RecognitionConfig\AudioEncoding::WEBM_OPUS, 'sample_rate_hertz' => 48000, 'language_code' => 'sk-SK', ]), 'interim_results' => true, ]); // 创建流式连接并发送配置 $stream = $speechClient->streamingRecognize(); $stream->write(new StreamingRecognizeRequest(['streaming_config' => $config])); // 异步读取转录结果,避免阻塞WebSocket连接 $this->listenForTranscripts($stream, $conn); // 保存资源,关联当前连接 $this->speechClients[$conn->resourceId] = $speechClient; $this->streams[$conn->resourceId] = $stream; } public function onMessage(ConnectionInterface $conn, $audioData) { // 将前端传来的音频块写入Google Speech流 if (isset($this->streams[$conn->resourceId])) { $this->streams[$conn->resourceId]->write(new StreamingRecognizeRequest([ 'audio_content' => $audioData ])); } } public function onClose(ConnectionInterface $conn) { // 关闭连接时清理资源,避免内存泄漏 if (isset($this->streams[$conn->resourceId])) { $this->streams[$conn->resourceId]->closeWrite(); unset($this->streams[$conn->resourceId]); } if (isset($this->speechClients[$conn->resourceId])) { $this->speechClients[$conn->resourceId]->close(); unset($this->speechClients[$conn->resourceId]); } } public function onError(ConnectionInterface $conn, \Exception $e) { $conn->close(); // 可添加错误日志记录 } protected function listenForTranscripts($stream, ConnectionInterface $conn) { // 异步处理转录结果,不阻塞WebSocket主线程 async(function () use ($stream, $conn) { while ($response = $stream->read()) { foreach ($response->getResults() as $result) { $transcript = $result->getAlternatives()[0]->getTranscript(); // 通过Pusher广播转录结果 event(new SpeechTranscriptCreated([ 'transcript' => $transcript, 'connection_id' => $conn->resourceId ])); } } }); } }
WebSocket路由配置
在routes/websockets.php中添加路由:
<?php use App\Http\Controllers\WebSocket\SpeechStreamController; Route::websocket('/ws/speech-stream', SpeechStreamController::class);
事件类优化(可选,精准广播)
修改SpeechTranscriptCreated事件,可指定只广播给当前用户:
<?php namespace App\Events; use Illuminate\Broadcasting\Channel; use Illuminate\Broadcasting\InteractsWithSockets; use Illuminate\Broadcasting\PrivateChannel; use Illuminate\Contracts\Broadcasting\ShouldBroadcast; use Illuminate\Foundation\Events\Dispatchable; use Illuminate\Queue\SerializesModels; class SpeechTranscriptCreated implements ShouldBroadcast { use Dispatchable, InteractsWithSockets, SerializesModels; public $transcript; public function __construct(array $data) { $this->transcript = $data['transcript']; } public function broadcastOn() { // 公共频道:所有订阅用户都能收到 return new Channel('speech-transcript-created-channel'); // 私有频道:仅当前认证用户收到(需结合Laravel Echo认证) // return new PrivateChannel('user.' . auth()->id()); } }
关键修复点说明
- WebSocket长连接:保证音频数据连续、实时地传输到后端,避免HTTP断流导致的Google Speech超时。
- 单连接对应单Speech流:每个WebSocket会话对应一个Google Speech流式连接,形成完整的音频会话。
- 启用临时结果:
interim_results: true让Google Speech返回实时的中间转录结果,满足实时翻译的低延迟需求。 - 资源清理:连接关闭时主动释放Google Speech的客户端和流资源,避免内存泄漏。
内容的提问来源于stack exchange,提问作者Peter
相关产品推荐
相关产品推荐

