如何将OpenAI API返回JSON中的带编号文本转为结构化列表
提取OpenAI API返回文本为结构化列表的解决方案
需求说明
从OpenAI API返回的JSON中提取choices[0].text字段内容,转换为无编号的字符串列表,要求:
- 移除开头的
\n\n1.类前缀 - 将
\n2.等带编号的换行替换为列表分隔 - 兼容API返回无编号、仅含
\n\n或\n的内容
示例API返回:
response = {'id': 'xyz', 'object': 'text_completion', 'created': 1673323957, 'model': 'text-davinci-003', 'choices': [{'text': '\n\n1. Dog Diet and Nutrition \n2. Dog Vaccination and Immunization \n3. Dog Parasites and Parasite Control \n4. Dog Dental Care and Hygiene \n5. Dog Grooming and Skin Care \n6. Dog Exercise and Training \n7. Dog First-Aid and Emergency Care \n8. Dog Joint Care and Arthritis \n9. Dog Allergies and Allergy Prevention \n10. Dog Senior Care and Health', 'index': 0, 'logprobs': None, 'finish_reason': 'length'}], 'usage': {'prompt_tokens': 16, 'completion_tokens': 100, 'total_tokens': 116}}
目标格式:
new_choices = ['Dog Diet and Nutrition', 'Dog Vaccination and Immunization', 'Dog Parasites and Parasite Control', 'Dog Dental Care and Hygiene', 'Dog Grooming and Skin Care', 'Dog Exercise and Training', 'Dog First-Aid and Emergency Care', 'Dog Joint Care and Arthritis', 'Dog Allergies and Allergy Prevention', 'Dog Senior Care and Health']
错误尝试
以下代码无法达到目标,结果保留编号且存在多余逗号:
new_choices = [response.json()['choices'][0]['text'].replace('\n',',')]
执行结果:
[',,1. Dog Diet and Nutrition ,2. Dog Vaccination and Immunization ,3. Dog Parasites and Parasite Control ,4. Dog Dental Care and Hygiene ,5. Dog Grooming and Skin Care ,6. Dog Exercise and Training ,7. Dog First-Aid and Emergency Care ,8. Dog Joint Care and Arthritis ,9. Dog Allergies and Allergy Prevention ,10. Dog Senior Care and Health']
正确解决方案
方案1:正则表达式处理(兼容多场景)
用正则匹配并移除编号,再拆分清理内容:
import re # 提取text内容(若为requests响应,需先调用response.json()解析) raw_text = response['choices'][0]['text'] # 移除所有数字编号(如1.、2.) cleaned_text = re.sub(r'\d+\.\s*', '', raw_text) # 按换行拆分,过滤空行并清理首尾空格 new_choices = [item.strip() for item in cleaned_text.split('\n') if item.strip()]
方案2:分步遍历处理(更直观)
无需正则,通过遍历行内容完成处理:
raw_text = response['choices'][0]['text'] # 按换行拆分所有行 lines = raw_text.split('\n') new_choices = [] for line in lines: stripped_line = line.strip() # 跳过空行 if not stripped_line: continue # 处理带编号的行:拆分编号与内容 if '.' in stripped_line[:3]: content = stripped_line.split('.', 1)[1].strip() else: content = stripped_line new_choices.append(content)
方案说明
- 两种方案均支持处理带编号的文本,自动移除
1.、2.类前缀 - 会过滤空行、清理条目首尾空格,兼容API返回仅含换行或无编号的内容
- 若使用
requests库调用API,需先通过.json()将响应解析为字典再提取字段
内容的提问来源于stack exchange,提问作者RustyShackleford
相关产品推荐
相关产品推荐

