Google Cloud STT代码报3 INVALID_ARGUMENT错误,求排查原因
Google Cloud Speech-to-Text v2 「Invalid resource field value」错误排查
问题场景
使用Node.js编写Google Cloud Speech-to-Text(STT)v2版本代码,传入的音频通过Buffer.from(someBase64EncodedAudioString, 'base64')转换为Buffer,已替换recognizer参数的项目占位符,但运行时抛出错误:
Error: 3 INVALID_ARGUMENT: Invalid resource field value in the request.
相关代码如下:
import { v2 } from '@google-cloud/speech'; import { fileURLToPath } from 'url'; import path from 'path'; const keyFilePath = './stt_key.json'; class STTService { constructor() { const __filename = fileURLToPath(import.meta.url); const __dirname = path.dirname(__filename); console.log(path.resolve(__dirname, keyFilePath)); this.client = new v2.SpeechClient({ keyFilename: path.resolve(__dirname, keyFilePath), }); this.openConnections = {}; this.request = { recognizer: 'projects/[project-name-here]/locations/global/recognizers/my-recognizer', streamingConfig: { config: { explicitDecodingConfig: { encoding: 'LINEAR16', sampleRateHertz: 16000, }, languageCodes: ['en-US'], model: 'latest_long', // or other model if applicable }, }, }; } async sendRequest(sessionId, audio, handler) { if (this.openConnections[sessionId]) { this.openConnections[sessionId].write(audio); return; } this.openConnections[sessionId] = this.client .streamingRecognize() .on('error', console.error) .on('data', (data) => { console.log(data); /* handler( data.results[0] && data.results[0].alternatives[0].transcript ); */ }); this.openConnections[sessionId].write(this.request); this.openConnections[sessionId].write(audio); } endRequest(sessionId) { this.openConnections[sessionId].end(); delete this.openConnections[sessionId]; } } const service = new STTService(); export default service;
排查方向及解决方案
1. Recognizer资源路径格式错误
STT v2要求recognizer路径必须严格遵循projects/{project-id}/locations/{location}/recognizers/{recognizer-id}格式,重点检查:
- 确认
[project-name-here]替换的是GCP项目ID(不是项目显示名称,项目ID是系统生成的唯一字符串,可在GCP控制台项目设置中查看) - 检查区域是否匹配:如果你的recognizer创建在非
global区域(比如us-central1),需将路径中的global改为对应区域 - 确认
my-recognizer是你实际创建的recognizer ID,可在GCP Speech-to-Text控制台的「Recognizers」页面查看详情
2. 请求结构与v2规范不匹配
STT v2的流式请求结构有明确要求,需确保request对象的嵌套层级正确,且recognizer与配置参数匹配:
- 确认
explicitDecodingConfig的编码、采样率与你创建recognizer时的设置一致 - 修正后的请求示例(确保结构正确):
this.request = { recognizer: 'projects/你的项目ID/locations/global/recognizers/my-recognizer', streamingConfig: { config: { explicitDecodingConfig: { encoding: 'LINEAR16', sampleRateHertz: 16000, }, languageCodes: ['en-US'], model: 'latest_long', }, // 如需返回实时中间结果,可添加该行 interimResults: true, }, };
3. 音频数据格式不匹配
虽然已将Base64转为Buffer,但需确认原始音频符合配置要求:
- 原始音频必须是裸PCM格式(无WAV文件头),编码为
LINEAR16,采样率16000Hz,单声道 - 如果原始音频是MP3、带文件头的WAV等格式,需先通过工具(如ffmpeg)解码为裸PCM:
ffmpeg -i input.mp3 -acodec pcm_s16le -ar 16000 -ac 1 -f s16le output.raw - 可通过
ffmpeg -i input.wav -f null -验证音频的采样率、编码参数
4. 权限或资源未正确配置
- 确认
stt_key.json对应的服务账号拥有speech.recognizers.use权限,可在GCP IAM页面为该账号添加「Speech-to-Text Recognizer User」角色 - 确认目标recognizer已在指定项目和区域中创建,且状态为正常(未被删除或禁用)
内容的提问来源于stack exchange,提问作者user28367041
相关产品推荐
相关产品推荐

