如何将Google Speech-To-Text API集成到网站实现语音填充textarea?
网页集成语音识别实现textarea填充方案
如果你需要在网页中实现语音输入填充textarea的功能,有两种可行方案,分别对应不同的需求场景:
方案一:使用浏览器原生Web Speech API(快速实现,无需谷歌云服务)
大部分现代浏览器(如Chrome)支持原生的Web Speech API,无需依赖谷歌云密钥,适合快速搭建基础功能。你提到的Demo正是基于该API实现的,之前可能混淆了Web Speech API和谷歌Cloud Speech-To-Text API的区别。
实现代码
HTML结构
<textarea id="speech-input" rows="5" cols="50" placeholder="点击开始按钮后说话,内容会自动填充这里"></textarea> <div> <button id="start-record">开始语音输入</button> <button id="stop-record" disabled>停止语音输入</button> </div>
JavaScript逻辑
// 检测浏览器是否支持语音识别 if ('webkitSpeechRecognition' in window) { const recognition = new webkitSpeechRecognition(); recognition.continuous = true; // 开启连续识别模式 recognition.interimResults = true; // 返回识别过程中的中间结果 recognition.lang = 'zh-CN'; // 设置识别语言,英文可改为'en-US' const textarea = document.getElementById('speech-input'); const startBtn = document.getElementById('start-record'); const stopBtn = document.getElementById('stop-record'); // 识别结果回调 recognition.onresult = function(event) { let fullTranscript = ''; // 遍历所有识别结果片段 for (let i = event.resultIndex; i < event.results.length; i++) { fullTranscript += event.results[i][0].transcript; } textarea.value = fullTranscript; }; // 识别开始时更新按钮状态 recognition.onstart = function() { startBtn.disabled = true; stopBtn.disabled = false; }; // 识别结束时恢复按钮状态 recognition.onend = function() { startBtn.disabled = false; stopBtn.disabled = true; }; // 绑定按钮点击事件 startBtn.addEventListener('click', () => recognition.start()); stopBtn.addEventListener('click', () => recognition.stop()); } else { alert('当前浏览器不支持语音识别功能,请使用Chrome浏览器尝试'); }
方案二:使用谷歌Cloud Speech-To-Text API(更高准确率,需云端服务)
如果需要更高的识别准确率或多语言支持,可以使用谷歌Cloud Speech-To-Text API,但该API是服务端接口,前端无法直接安全调用(会暴露密钥),需通过后端代理实现。
实现步骤
- 在谷歌云平台创建项目,启用Speech-To-Text API,生成服务账号密钥文件。
- 搭建后端服务代理API请求,前端录制音频后发送到后端,后端调用云端API返回结果。
前端代码(录制音频并发送)
let mediaRecorder; let audioChunks = []; const textarea = document.getElementById('speech-input'); async function startCloudRecording() { try { const stream = await navigator.mediaDevices.getUserMedia({ audio: true }); mediaRecorder = new MediaRecorder(stream); audioChunks = []; mediaRecorder.ondataavailable = (event) => { audioChunks.push(event.data); }; mediaRecorder.onstop = async () => { const audioBlob = new Blob(audioChunks, { type: 'audio/wav' }); const formData = new FormData(); formData.append('audio', audioBlob, 'recording.wav'); // 发送到后端代理接口 const response = await fetch('/api/recognize-speech', { method: 'POST', body: formData }); const result = await response.json(); textarea.value = result.transcript; }; mediaRecorder.start(); document.getElementById('cloud-start').disabled = true; document.getElementById('cloud-stop').disabled = false; } catch (err) { console.error('录制启动失败:', err); } } function stopCloudRecording() { mediaRecorder.stop(); document.getElementById('cloud-start').disabled = false; document.getElementById('cloud-stop').disabled = true; }
后端代码(Node.js示例)
const express = require('express'); const multer = require('multer'); const speech = require('@google-cloud/speech'); const fs = require('fs'); const app = express(); const upload = multer({ dest: 'temp-uploads/' }); // 初始化谷歌云Speech客户端 const client = new speech.SpeechClient({ keyFilename: './service-account-key.json' // 替换为你的服务账号密钥路径 }); // 语音识别接口 app.post('/api/recognize-speech', upload.single('audio'), async (req, res) => { try { const audioContent = fs.readFileSync(req.file.path).toString('base64'); const request = { audio: { content: audioContent }, config: { encoding: 'LINEAR16', sampleRateHertz: 16000, languageCode: 'zh-CN' } }; const [response] = await client.recognize(request); const transcript = response.results .map(result => result.alternatives[0].transcript) .join('\n'); res.json({ transcript }); // 删除临时音频文件 fs.unlinkSync(req.file.path); } catch (err) { console.error('识别失败:', err); res.status(500).json({ error: '语音识别失败' }); } }); app.listen(3000, () => console.log('后端服务运行在端口3000'));
你之前代码的问题说明
你之前使用gapi初始化谷歌Cloud API的方式存在三个核心问题:
- CORS限制:谷歌Cloud API不允许前端直接跨域调用,会被浏览器拦截。
- 安全风险:前端暴露API密钥或Client ID会导致密钥泄露,被恶意使用。
- 缺少音频处理逻辑:没有实现音频录制、格式转换和API请求发送的完整流程,因此无法获取识别结果。
如果只是需要基础的语音填充功能,方案一的Web Speech API完全可以满足需求,且无需额外配置云端服务。
内容的提问来源于stack exchange,提问作者janco
相关产品推荐
相关产品推荐

