如何在使用OpenAI GPT-3 API时保留响应格式?
解决GPT-3 API返回文本格式丢失的问题
1. 明确Prompt中的格式要求
调用API时,在prompt里直接指定输出格式规则,让模型生成带换行、列表结构的文本。比如在prompt末尾补充:
请以清晰的格式输出内容:段落与列表之间留空行,编号列表用「1. 」开头,无序列表用「- 」开头,每个列表项单独占一行。
也可以在prompt里加入Playground的格式示例,引导模型参考生成。
2. 对返回文本做后处理修复
如果模型返回的文本仍为紧凑格式,可通过字符串处理工具修复:
- 编号列表修复:用正则表达式在每个
数字.前插入换行(排除文本开头的情况),比如在Python中用re.sub(r'(?<!^)(\d+\.)', r'\n\1', raw_text)。 - 无序列表修复:在每个
-前插入换行,用re.sub(r'(- )', r'\n\1', raw_text)处理。 - 衔接处修复:针对类似“the following”这类衔接文本,在其后插入换行,避免列表直接粘连,比如
re.sub(r'(the following)(-)', r'\1\n\2', raw_text)。
示例Python代码:
import re def fix_gpt_format(raw_text): # 修复编号列表换行 formatted = re.sub(r'(?<!^)(\d+\.)', r'\n\1', raw_text) # 修复无序列表换行 formatted = re.sub(r'(- )', r'\n\1', formatted) # 修复标题与编号的衔接 formatted = re.sub(r'(doing:)(\d+\.)', r'\1\n\2', formatted) # 修复"the following"与列表的衔接 formatted = re.sub(r'(the following)(-)', r'\1\n\2', formatted) return formatted # 测试示例 raw_content = "Here's what the above class is doing:1. It creates a directory for the log file if it doesn't exist.2. It checks that the log file is newline-terminated.3. It writes a newline-terminated JSON object to the log file.4. It reads the log file and returns a dictionary with the following-list 1-list 2-list 3- list4" print(fix_gpt_format(raw_content))
3. 切换至更优模型
GPT-4及后续迭代模型对格式指令的遵循度远高于GPT-3,能更精准地生成符合要求的结构化文本,大幅减少后处理的工作量。
内容的提问来源于stack exchange,提问作者Tyler Kim
相关产品推荐
相关产品推荐

