如何获取GitHub Copilot Chat Completions API的LLM响应文本
如何从GitHub Copilot Chat Completions API的流式SSE响应中提取纯文本内容
问题描述
我正在开发GitHub Copilot扩展,基于Copilot LLM的Chat Completions API实现功能。之前参考示例代码能通过pipe直接将流式响应输出给用户,但现在需要自行提取响应中的纯文本内容。
当前可正常请求的代码:
// Use Copilot's LLM to generate a response to the user's messages, with // our extra system messages attached. const copilotLLMResponse = await fetch( "https://api.githubcopilot.com/chat/completions", { method: "POST", headers: { authorization: `Bearer ${tokenForUser}`, "content-type": "application/json", }, body: JSON.stringify({ messages, stream: true, }), } ); const stream = Readable.from(copilotLLMResponse.body); stream.pipe(res);
尝试按照Bun的流转字符串指南改写后,仅能获取带data:前缀的SSE格式内容,无法提取纯文本:
data: {"choices":[{"finish_reason":"stop","index":0,"content_filter_offsets":{"check_offset":1896,"start_offset":1792,"end_offset":1897},"content_filter_results":{"error":{"code":"","message":""},"hate":{"filtered":false,"severity":"safe"},"self_harm":{"filtered":false,"severity":"safe"},"sexual":{"filtered":false,"severity":"safe"},"violence":{"filtered":false,"severity":"safe"}},"delta":{"content":null}}],"created":1730025876,"id":"chatcmpl-someId","usage":{"completion_tokens":22,"prompt_tokens":404,"total_tokens":426},"model":"gpt-3.5-turbo-0613"}
我的最终需求是向Copilot注入自定义文档,使其生成符合要求的用户响应,因此必须先解决纯文本提取问题。
解决方案
GitHub Copilot的流式响应采用Server-Sent Events(SSE)格式,每个数据块以data:开头,后续为JSON结构。要提取纯文本,需逐行解析这些数据块,过滤无效内容后提取delta.content字段。
以下是适配Bun环境的实现代码:
const copilotLLMResponse = await fetch( "https://api.githubcopilot.com/chat/completions", { method: "POST", headers: { authorization: `Bearer ${tokenForUser}`, "content-type": "application/json", }, body: JSON.stringify({ messages, stream: true, }), } ); const reader = copilotLLMResponse.body.getReader(); const decoder = new TextDecoder(); let fullResponse = ''; while (true) { const { done, value } = await reader.read(); if (done) break; // 解码二进制数据为文本 const chunk = decoder.decode(value, { stream: true }); // 按行分割并过滤空行 const lines = chunk.split('\n').filter(line => line.trim() !== ''); for (const line of lines) { if (line.startsWith('data:')) { const jsonStr = line.slice(5).trim(); // 跳过流式结束标记 if (jsonStr === '[DONE]') continue; try { const data = JSON.parse(jsonStr); // 提取当前片段的文本内容 const content = data.choices?.[0]?.delta?.content; if (content) { fullResponse += content; // 如需实时输出,可在此处直接处理单段内容,例如: // process.stdout.write(content); } } catch (err) { console.error('解析SSE数据失败:', err); } } } } // 最终提取的完整纯文本响应 console.log('完整响应:', fullResponse);
核心逻辑说明
- 读取流式数据:使用
getReader()逐块读取响应流的二进制数据 - 解码文本:通过
TextDecoder将二进制数据转换为可读字符串 - 解析SSE格式:按行分割数据,仅处理以
data:开头的行,移除前缀后解析JSON - 提取有效内容:从JSON结构中获取
choices[0].delta.content,这就是流式返回的纯文本片段 - 实时/批量处理:若需实时输出给用户,可在获取到单段
content时直接处理;若需完整响应,可将片段拼接为完整字符串
内容的提问来源于stack exchange,提问作者Anay
相关产品推荐
相关产品推荐

