You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Symbl Streaming API实时转录缺失user/from字段问题咨询

问题

已通过WebSocket连接Symbl Streaming API,发送的start_request请求如下:

{
  type: "start_request",
  meetingTitle: "Symbl Connect SIP",
  config: {
    confidenceThreshold: 0.5,
    languageCode: "en-US",
    speechRecognition: {
      encoding: "LINEAR16",
      sampleRateHertz: userType === "user" ? this.sampleRateUser : this.sampleRateAgent,
    },
    trackers: {
      enableAllTrackers: true,
      interimResults: true,
    },
    speaker: {
      ...speaker,
      name: userType,
    },
  },
}

绑定onmessage事件后,返回的实时转录(recognition_result)和message_response数据中缺失包含名称及userId的user或from字段:

  • message_response 实际输出
messages: [
  {
    payload: {
      content: string,
      contentType: "text/plain",
    },
    id: string,
    channel: { id: "realtime-api" },
    metadata: {
      disablePunctuation: true,
      originalContent: string,
      words: string,
      originalMessageId: string,
    },
    dismissed: false,
    duration: {
      startTime: string,
      endTime: string,
      timeOffset: number,
      duration: number,
    },
    entities: [],
  },
]
  • 实时转录 实际输出
{
  type: "recognition_result",
  isFinal: false,
  payload: {
    raw: {
      alternatives: [
        {
          words: [],
          transcript: "Hello",
          confidence: 0,
        },
      ],
    },
  },
  punctuated: { transcript: "Hello" },
}

期望返回结果包含user(转录结果)或from(message_response)字段,示例如下:

  • 期望的实时转录输出
{
  type: "recognition_result",
  isFinal: false,
  payload: {
    raw: {
      alternatives: [
        {
          words: [],
          transcript: "Hello",
          confidence: 0,
        },
      ],
    },
  },
  punctuated: { transcript: "Hello" },
  user: {
    name: "agent",
    userId: "agent@123.com",
    id: "1a726dac-1c95-484d-a99e-69e36950bf32"
  }
}
  • 期望的message_response输出
{
  from: {
    id: string,
    name: string,
    userId: string,
  },
  payload: {
    content: string,
    contentType: "text/plain",
  },
  id: string,
  channel: { id: "realtime-api" },
  metadata: {
    disablePunctuation: true,
    originalContent: string,
    words: string,
    originalMessageId: string,
  },
  dismissed: false,
  duration: {
    startTime: string,
    endTime: string,
    timeOffset: number,
    duration: number,
  },
  entities: [],
}

请问该如何配置才能让返回结果包含对应的user/from字段?

解决方法

问题出在start_request的配置结构和字段完整性上,需要做两处修改:

  1. 调整speaker字段的层级
    当前把speaker放在了config对象内部,但Symbl的start_request规范中,speaker是和config、meetingTitle同级的顶层字段,不属于config的子属性。

  2. 补充speaker的完整字段
    要让返回结果带上用户信息,speaker对象需要包含id、name(必填),userId(可选但建议补充),这些字段会映射到返回结果的user/from属性中。

修改后的start_request示例:

{
  type: "start_request",
  meetingTitle: "Symbl Connect SIP",
  config: {
    confidenceThreshold: 0.5,
    languageCode: "en-US",
    speechRecognition: {
      encoding: "LINEAR16",
      sampleRateHertz: userType === "user" ? this.sampleRateUser : this.sampleRateAgent,
    },
    trackers: {
      enableAllTrackers: true,
      interimResults: true,
    },
  },
  // speaker移到顶层,补充完整标识字段
  speaker: {
    id: userType === "user" ? "user-123" : "agent-456", // 自定义唯一标识
    name: userType,
    userId: userType === "user" ? "user@example.com" : "agent@example.com"
  }
}

如果是多说话人场景,后续发送音频数据时,可通过audio_request的speaker字段指定当前音频对应的说话人(与start_request中配置不同时),示例:

{
  type: "audio_request",
  audio: <base64-encoded-audio>,
  speaker: {
    id: "guest-789",
    name: "Guest",
    userId: "guest@example.com"
  }
}

配置完成后,实时转录的recognition_result会带上user字段,message_response中的每条消息会带上from字段,包含你设置的说话人信息。

内容的提问来源于stack exchange,提问作者Boanerges

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 05:05:16