如何避免ChatGPT Realtime单响应中语音朗读JSON内容?
使用ChatGPT Realtime API实现语音命令识别的问题与解决方案
需求与问题
我正在基于ChatGPT Realtime API开发语音命令识别与自然语言对话功能,核心需求如下:
- 用户说出命令(如「打开客厅灯」)时,API需返回两项内容:
- 结构化文本输出(如JSON,含命令ID,无命令则为空),供程序解析执行;
- 自然语音输出(如「好的,正在开灯」),而非朗读原始JSON。
- 用户表述非命令时,API仅需返回自然语音回应即可。
当前遇到的核心问题:启用音频输出模态(modality)时,模型会朗读所有文本输出内容,导致语音会读出JSON结构。
相关代码
// Node 18+ // npm i ws speaker import WebSocket from 'ws'; import fs from 'fs'; import Speaker from 'speaker'; const OPENAI_API_KEY = process.env.OPENAI_API_KEY; if (!OPENAI_API_KEY) { console.error('Set OPENAI_API_KEY'); process.exit(1); } const REALTIME_URL = 'wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview'; const SAMPLE_RATE = 24000; const MODE = process.env.MODE || 'BAD'; // BAD | GOOD const PCM_FILE = process.env.PCM_FILE || 'sample.pcm'; // Basic audio out (PCM16 mono 24kHz) const speaker = new Speaker({ channels: 1, bitDepth: 16, sampleRate: SAMPLE_RATE, signed: true }); function playPcmBase64Chunk(b64) { const buf = Buffer.from(b64, 'base64'); speaker.write(buf); } const ws = new WebSocket(REALTIME_URL, { headers: { Authorization: `Bearer ${OPENAI_API_KEY}`, 'OpenAI-Beta': 'realtime=v1' } }); // Buffers to capture outputs for debugging let textBuf = ''; let audioStarted = false; ws.on('open', async () => { console.log('[WS] connected, setting up session…'); // Session with audio in/out PCM16. (No tool calling; only text+audio.) ws.send(JSON.stringify({ type: 'session.update', session: { modalities: ['text', 'audio'], // session supports both, but we’ll control per-response input_audio_format: { type: 'pcm16', sample_rate: SAMPLE_RATE }, output_audio_format: { type: 'pcm16', sample_rate: SAMPLE_RATE } } })); // Feed a short PCM16 sample (1–2 seconds) const pcm = fs.readFileSync(PCM_FILE); const CHUNK = 8000; // arbitrary small chunks for (let i = 0; i < pcm.length; i += CHUNK) { const slice = pcm.subarray(i, i + CHUNK); ws.send(JSON.stringify({ type: 'input_audio_buffer.append', audio: slice.toString('base64') })); } // finalize audio input ws.send(JSON.stringify({ type: 'input_audio_buffer.commit' })); if (MODE === 'BAD') { console.log('\nMODE=BAD → one response with ["text","audio"]'); ws.send(JSON.stringify({ type: 'response.create', response: { // both outputs in ONE response: modalities: ['text', 'audio'], // ask for JSON in text and a natural confirmation in audio // but sometimes the audio ends up speaking JSON :( instructions: 'Return ONLY JSON in the TEXT modality (machine-readable command). ' + 'Then produce a short, natural confirmation in the AUDIO modality. ' + 'Do NOT read or mention any JSON in speech.' } })); } else { console.log('\nMODE=GOOD → two responses: [text-only] then [audio-only]'); // 1) TEXT-ONLY: get JSON (works reliably) ws.send(JSON.stringify({ type: 'response.create', response: { modalities: ['text'], instructions: 'Return ONLY JSON. No extra words. JSON must represent the user command (intent, device, location, value).' } })); // 2) AUDIO-ONLY: speak a natural confirmation (no JSON spoken) // Small delay to avoid races; ideally wait for response.done of the previous. setTimeout(() => { ws.send(JSON.stringify({ type: 'response.create', response: { modalities: ['audio'], instructions: 'Say a short, friendly confirmation. Do NOT read or mention any JSON.' } })); }, 600); } }); // Handle server events (text + audio) ws.on('message', (data) => { const msg = JSON.parse(data); // Text streaming if (msg.type === 'response.output_text.delta' && msg.delta) { textBuf += msg.delta; } if (msg.type === 'response.output_text.done') { console.log('\n[TEXT DONE]\n' + textBuf); } // Audio streaming if (msg.type === 'response.output_audio.delta' && msg.delta) { if (!audioStarted) { audioStarted = true; console.log('\n[AUDIO streaming…]'); } playPcmBase64Chunk(msg.delta); } if (msg.type === 'response.output_audio.done') { console.log('\n[AUDIO DONE]'); try { speaker.end(); } catch {} } // For debugging if (msg.type === 'error') { console.error('[API ERROR]', msg); } }); process.on('SIGINT', () => { try { speaker.end(); } catch {} try { ws.close(); } catch {} process.exit(0); });
当前两种模式表现
- BAD模式:单次调用
response.create同时启用text和audio模态,期望文本返回纯JSON、语音为自然确认,但实际语音偶尔会朗读JSON内容。 - GOOD模式:分两次调用
response.create(先请求仅返回JSON的文本响应,再请求仅返回语音确认的音频响应),功能正常但会额外消耗输入token。
问题解答
1. 单次response.create同时使用text和audio模态时,如何确保语音绝不朗读JSON?
可以通过系统提示词全局约束+明确的模态任务拆分指令实现,核心是让模型清晰认知两个模态的独立任务边界:
- 在会话初始化时添加
system_prompt,提前告知模型两个模态的严格分工; - 在
response.create的指令中再次强化,明确音频模态完全忽略文本输出内容。
这种方式能大幅降低模型混淆两个任务的概率,确保语音仅输出自然回应。
2. 有没有推荐的提示词模式或API标记来标记非语音合成内容?
目前Realtime API没有内置的标记来屏蔽文本的语音合成,但可以通过强约束的结构化提示词实现:
- 采用「任务拆分+明确禁止」的提示词结构,比如:
【文本任务】仅输出机器可读的JSON,无任何额外内容;【语音任务】仅输出面向用户的自然口语回应,完全忽略文本任务的内容,绝不提及或朗读JSON。 - 在会话初始化时设置全局系统提示,示例如下:
{ "type": "session.update", "session": { "modalities": ["text", "audio"], "input_audio_format": {"type": "pcm16", "sample_rate": 24000}, "output_audio_format": {"type": "pcm16", "sample_rate": 24000}, "system_prompt": "你需要完成两个独立的输出任务,严格区分模态:\n1. 文本模态:仅返回结构化JSON,对应用户的命令意图(无命令则返回{}),无任何额外文字;\n2. 音频模态:仅返回自然、简洁的用户语音回应,完全忽略文本模态的内容,绝不提及或朗读JSON。" } }
3. 可用示例与最佳实践
优化后的单次response.create示例
修改BAD模式的会话初始化和响应请求,强化约束:
ws.on('open', async () => { console.log('[WS] connected, setting up session…'); // 会话初始化添加全局系统提示 ws.send(JSON.stringify({ type: 'session.update', session: { modalities: ['text', 'audio'], input_audio_format: { type: 'pcm16', sample_rate: SAMPLE_RATE }, output_audio_format: { type: 'pcm16', sample_rate: SAMPLE_RATE }, system_prompt: "你需要完成两个独立的输出任务,严格区分模态:\n1. 文本模态:仅返回结构化JSON,对应用户的命令意图(无命令则返回{}),无任何额外文字;\n2. 音频模态:仅返回自然、简洁的用户语音回应,完全忽略文本模态的内容,绝不提及或朗读JSON。" } })); // 音频输入逻辑不变... // 单次response.create调用 console.log('\nMODE=OPTIMIZED → one response with ["text","audio"]'); ws.send(JSON.stringify({ type: 'response.create', response: { modalities: ['text', 'audio'], instructions: "文本模态输出对应用户命令的JSON;音频模态输出自然的口语确认,比如用户说'打开客厅灯'就回应'好的,正在打开客厅灯',不要涉及任何JSON相关内容。" } })); });
最佳实践
- 优先全局系统提示:比单次响应指令更稳定,模型会全程遵守规则;
- 指令要绝对明确:避免模糊表述,明确告知模型两个模态是完全独立的任务,音频模态无需关注文本输出;
- 测试边缘场景:比如用户输入非命令内容(如「今天天气怎么样」),验证文本返回空JSON、语音正常回答;
- 增加约束强度:如果仍出现JSON朗读,可在提示词中添加惩罚性描述:
如果音频模态提及或朗读JSON,视为错误,必须重新生成纯自然语音回应。
内容的提问来源于stack exchange,提问作者danflu
相关产品推荐
相关产品推荐

