Symbl Streaming API实时转录缺失user/from字段问题咨询
问题
已通过WebSocket连接Symbl Streaming API,发送的start_request请求如下:
{ type: "start_request", meetingTitle: "Symbl Connect SIP", config: { confidenceThreshold: 0.5, languageCode: "en-US", speechRecognition: { encoding: "LINEAR16", sampleRateHertz: userType === "user" ? this.sampleRateUser : this.sampleRateAgent, }, trackers: { enableAllTrackers: true, interimResults: true, }, speaker: { ...speaker, name: userType, }, }, }
绑定onmessage事件后,返回的实时转录(recognition_result)和message_response数据中缺失包含名称及userId的user或from字段:
- message_response 实际输出
messages: [ { payload: { content: string, contentType: "text/plain", }, id: string, channel: { id: "realtime-api" }, metadata: { disablePunctuation: true, originalContent: string, words: string, originalMessageId: string, }, dismissed: false, duration: { startTime: string, endTime: string, timeOffset: number, duration: number, }, entities: [], }, ]
- 实时转录 实际输出
{ type: "recognition_result", isFinal: false, payload: { raw: { alternatives: [ { words: [], transcript: "Hello", confidence: 0, }, ], }, }, punctuated: { transcript: "Hello" }, }
期望返回结果包含user(转录结果)或from(message_response)字段,示例如下:
- 期望的实时转录输出
{ type: "recognition_result", isFinal: false, payload: { raw: { alternatives: [ { words: [], transcript: "Hello", confidence: 0, }, ], }, }, punctuated: { transcript: "Hello" }, user: { name: "agent", userId: "agent@123.com", id: "1a726dac-1c95-484d-a99e-69e36950bf32" } }
- 期望的message_response输出
{ from: { id: string, name: string, userId: string, }, payload: { content: string, contentType: "text/plain", }, id: string, channel: { id: "realtime-api" }, metadata: { disablePunctuation: true, originalContent: string, words: string, originalMessageId: string, }, dismissed: false, duration: { startTime: string, endTime: string, timeOffset: number, duration: number, }, entities: [], }
请问该如何配置才能让返回结果包含对应的user/from字段?
解决方法
问题出在start_request的配置结构和字段完整性上,需要做两处修改:
调整speaker字段的层级
当前把speaker放在了config对象内部,但Symbl的start_request规范中,speaker是和config、meetingTitle同级的顶层字段,不属于config的子属性。补充speaker的完整字段
要让返回结果带上用户信息,speaker对象需要包含id、name(必填),userId(可选但建议补充),这些字段会映射到返回结果的user/from属性中。
修改后的start_request示例:
{ type: "start_request", meetingTitle: "Symbl Connect SIP", config: { confidenceThreshold: 0.5, languageCode: "en-US", speechRecognition: { encoding: "LINEAR16", sampleRateHertz: userType === "user" ? this.sampleRateUser : this.sampleRateAgent, }, trackers: { enableAllTrackers: true, interimResults: true, }, }, // speaker移到顶层,补充完整标识字段 speaker: { id: userType === "user" ? "user-123" : "agent-456", // 自定义唯一标识 name: userType, userId: userType === "user" ? "user@example.com" : "agent@example.com" } }
如果是多说话人场景,后续发送音频数据时,可通过audio_request的speaker字段指定当前音频对应的说话人(与start_request中配置不同时),示例:
{ type: "audio_request", audio: <base64-encoded-audio>, speaker: { id: "guest-789", name: "Guest", userId: "guest@example.com" } }
配置完成后,实时转录的recognition_result会带上user字段,message_response中的每条消息会带上from字段,包含你设置的说话人信息。
内容的提问来源于stack exchange,提问作者Boanerges

