You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免ChatGPT Realtime单响应中语音朗读JSON内容?

使用ChatGPT Realtime API实现语音命令识别的问题与解决方案

需求与问题

我正在基于ChatGPT Realtime API开发语音命令识别与自然语言对话功能,核心需求如下:

  • 用户说出命令(如「打开客厅灯」)时,API需返回两项内容:
    1. 结构化文本输出(如JSON,含命令ID,无命令则为空),供程序解析执行;
    2. 自然语音输出(如「好的,正在开灯」),而非朗读原始JSON。
  • 用户表述非命令时,API仅需返回自然语音回应即可。

当前遇到的核心问题:启用音频输出模态(modality)时,模型会朗读所有文本输出内容,导致语音会读出JSON结构。

相关代码

// Node 18+
// npm i ws speaker
import WebSocket from 'ws';
import fs from 'fs';
import Speaker from 'speaker';

const OPENAI_API_KEY = process.env.OPENAI_API_KEY;
if (!OPENAI_API_KEY) {
  console.error('Set OPENAI_API_KEY'); process.exit(1);
}

const REALTIME_URL = 'wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview';
const SAMPLE_RATE = 24000;
const MODE = process.env.MODE || 'BAD'; // BAD | GOOD
const PCM_FILE = process.env.PCM_FILE || 'sample.pcm';

// Basic audio out (PCM16 mono 24kHz)
const speaker = new Speaker({ channels: 1, bitDepth: 16, sampleRate: SAMPLE_RATE, signed: true });
function playPcmBase64Chunk(b64) {
  const buf = Buffer.from(b64, 'base64');
  speaker.write(buf);
}

const ws = new WebSocket(REALTIME_URL, {
  headers: {
    Authorization: `Bearer ${OPENAI_API_KEY}`,
    'OpenAI-Beta': 'realtime=v1'
  }
});

// Buffers to capture outputs for debugging
let textBuf = '';
let audioStarted = false;

ws.on('open', async () => {
  console.log('[WS] connected, setting up session…');

  // Session with audio in/out PCM16. (No tool calling; only text+audio.)
  ws.send(JSON.stringify({
    type: 'session.update',
    session: {
      modalities: ['text', 'audio'], // session supports both, but we’ll control per-response
      input_audio_format:  { type: 'pcm16', sample_rate: SAMPLE_RATE },
      output_audio_format: { type: 'pcm16', sample_rate: SAMPLE_RATE }
    }
  }));

  // Feed a short PCM16 sample (1–2 seconds)
  const pcm = fs.readFileSync(PCM_FILE);
  const CHUNK = 8000; // arbitrary small chunks
  for (let i = 0; i < pcm.length; i += CHUNK) {
    const slice = pcm.subarray(i, i + CHUNK);
    ws.send(JSON.stringify({
      type: 'input_audio_buffer.append',
      audio: slice.toString('base64')
    }));
  }
  // finalize audio input
  ws.send(JSON.stringify({ type: 'input_audio_buffer.commit' }));

  if (MODE === 'BAD') {
    console.log('\nMODE=BAD → one response with ["text","audio"]');
    ws.send(JSON.stringify({
      type: 'response.create',
      response: {
        // both outputs in ONE response:
        modalities: ['text', 'audio'],
        // ask for JSON in text and a natural confirmation in audio
        // but sometimes the audio ends up speaking JSON :(
        instructions:
          'Return ONLY JSON in the TEXT modality (machine-readable command). ' +
          'Then produce a short, natural confirmation in the AUDIO modality. ' +
          'Do NOT read or mention any JSON in speech.'
      }
    }));
  } else {
    console.log('\nMODE=GOOD → two responses: [text-only] then [audio-only]');
    // 1) TEXT-ONLY: get JSON (works reliably)
    ws.send(JSON.stringify({
      type: 'response.create',
      response: {
        modalities: ['text'],
        instructions:
          'Return ONLY JSON. No extra words. JSON must represent the user command (intent, device, location, value).'
      }
    }));

    // 2) AUDIO-ONLY: speak a natural confirmation (no JSON spoken)
    // Small delay to avoid races; ideally wait for response.done of the previous.
    setTimeout(() => {
      ws.send(JSON.stringify({
        type: 'response.create',
        response: {
          modalities: ['audio'],
          instructions:
            'Say a short, friendly confirmation. Do NOT read or mention any JSON.'
        }
      }));
    }, 600);
  }
});

// Handle server events (text + audio)
ws.on('message', (data) => {
  const msg = JSON.parse(data);

  // Text streaming
  if (msg.type === 'response.output_text.delta' && msg.delta) {
    textBuf += msg.delta;
  }
  if (msg.type === 'response.output_text.done') {
    console.log('\n[TEXT DONE]\n' + textBuf);
  }

  // Audio streaming
  if (msg.type === 'response.output_audio.delta' && msg.delta) {
    if (!audioStarted) { audioStarted = true; console.log('\n[AUDIO streaming…]'); }
    playPcmBase64Chunk(msg.delta);
  }
  if (msg.type === 'response.output_audio.done') {
    console.log('\n[AUDIO DONE]');
    try { speaker.end(); } catch {}
  }

  // For debugging
  if (msg.type === 'error') {
    console.error('[API ERROR]', msg);
  }
});

process.on('SIGINT', () => { try { speaker.end(); } catch {} try { ws.close(); } catch {} process.exit(0); });

当前两种模式表现

  • BAD模式:单次调用response.create同时启用text和audio模态,期望文本返回纯JSON、语音为自然确认,但实际语音偶尔会朗读JSON内容。
  • GOOD模式:分两次调用response.create(先请求仅返回JSON的文本响应,再请求仅返回语音确认的音频响应),功能正常但会额外消耗输入token。

问题解答

1. 单次response.create同时使用text和audio模态时,如何确保语音绝不朗读JSON?

可以通过系统提示词全局约束+明确的模态任务拆分指令实现,核心是让模型清晰认知两个模态的独立任务边界:

  1. 在会话初始化时添加system_prompt,提前告知模型两个模态的严格分工;
  2. 在response.create的指令中再次强化,明确音频模态完全忽略文本输出内容。

这种方式能大幅降低模型混淆两个任务的概率,确保语音仅输出自然回应。

2. 有没有推荐的提示词模式或API标记来标记非语音合成内容?

目前Realtime API没有内置的标记来屏蔽文本的语音合成,但可以通过强约束的结构化提示词实现:

  • 采用「任务拆分+明确禁止」的提示词结构,比如:
    【文本任务】仅输出机器可读的JSON,无任何额外内容;【语音任务】仅输出面向用户的自然口语回应,完全忽略文本任务的内容,绝不提及或朗读JSON。
  • 在会话初始化时设置全局系统提示,示例如下:
    {
      "type": "session.update",
      "session": {
        "modalities": ["text", "audio"],
        "input_audio_format": {"type": "pcm16", "sample_rate": 24000},
        "output_audio_format": {"type": "pcm16", "sample_rate": 24000},
        "system_prompt": "你需要完成两个独立的输出任务,严格区分模态:\n1. 文本模态:仅返回结构化JSON,对应用户的命令意图(无命令则返回{}),无任何额外文字;\n2. 音频模态:仅返回自然、简洁的用户语音回应,完全忽略文本模态的内容,绝不提及或朗读JSON。"
      }
    }
    

3. 可用示例与最佳实践

优化后的单次response.create示例

修改BAD模式的会话初始化和响应请求,强化约束:

ws.on('open', async () => {
  console.log('[WS] connected, setting up session…');

  // 会话初始化添加全局系统提示
  ws.send(JSON.stringify({
    type: 'session.update',
    session: {
      modalities: ['text', 'audio'],
      input_audio_format:  { type: 'pcm16', sample_rate: SAMPLE_RATE },
      output_audio_format: { type: 'pcm16', sample_rate: SAMPLE_RATE },
      system_prompt: "你需要完成两个独立的输出任务,严格区分模态:\n1. 文本模态:仅返回结构化JSON,对应用户的命令意图(无命令则返回{}),无任何额外文字;\n2. 音频模态:仅返回自然、简洁的用户语音回应,完全忽略文本模态的内容,绝不提及或朗读JSON。"
    }
  }));

  // 音频输入逻辑不变...

  // 单次response.create调用
  console.log('\nMODE=OPTIMIZED → one response with ["text","audio"]');
  ws.send(JSON.stringify({
    type: 'response.create',
    response: {
      modalities: ['text', 'audio'],
      instructions: "文本模态输出对应用户命令的JSON;音频模态输出自然的口语确认,比如用户说'打开客厅灯'就回应'好的,正在打开客厅灯',不要涉及任何JSON相关内容。"
    }
  }));
});

最佳实践

  • 优先全局系统提示:比单次响应指令更稳定,模型会全程遵守规则;
  • 指令要绝对明确:避免模糊表述,明确告知模型两个模态是完全独立的任务,音频模态无需关注文本输出;
  • 测试边缘场景:比如用户输入非命令内容(如「今天天气怎么样」),验证文本返回空JSON、语音正常回答;
  • 增加约束强度:如果仍出现JSON朗读,可在提示词中添加惩罚性描述:如果音频模态提及或朗读JSON,视为错误,必须重新生成纯自然语音回应。

内容的提问来源于stack exchange,提问作者danflu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 07:55:58