求助:修复LLM项目中Python JSON输出格式化问题
问题:LLM反馈文本转JSON格式异常修复
问题描述
开发LLM项目时,需要将LLM生成的反馈响应格式化为指定结构的JSON。现有Python代码通过正则表达式提取标题(issue)、node_ids、详细描述(detailed_feedback),但输出格式不符合预期,无法正确关联每个问题的对应字段。
原代码
import re import os import json from nltk.tokenize import sent_tokenize def extract_data(text): section_pattern = r'\d+\.\s\*(.*?)\((Node ID:.*?)\).*?((?=\d+\.\s*\*)|$)|\-\s(.*?)\n\n' section_regex = re.compile(section_pattern, re.MULTILINE | re.DOTALL) matches = section_regex.findall(text) data = [] # print(matches) for match in matches: heading = match[0].strip() node_ids = re.findall(r'\d+:\d+', match[1]) detailed_desc = extract_detailed_desc(match[3].strip()) data.append({ 'issue': heading, 'node_ids': node_ids, 'detailed_feedback': detailed_desc }) return data def extract_detailed_desc(text): sentences = sent_tokenize(text) detailed_desc = [] for sentence in sentences: detailed_desc.append(sentence.strip('-').strip()) return detailed_desc def main(): txt_file_path = "./FormatOutput/sample.txt" if os.path.exists(txt_file_path): try: with open(txt_file_path, 'r') as txt_file: data = txt_file.read() # print(data) except Exception as e: print("Error occurred while reading the text file:", e) else: print("File not found:", txt_file_path) structured_data = extract_data(data) json_data = json.dumps(structured_data, indent=4) print(json_data) if __name__ == "__main__": main()
问题分析
- 正则表达式采用分支匹配,将问题标题和详情项分开捕获,导致标题与详情无法正确关联,出现字段缺失或错位。
- 使用
sent_tokenize拆分详情文本,会把单个列表项拆分成多个句子,不符合预期的详情列表结构。 - 正则的边界匹配逻辑不严谨,无法完整捕获每个问题区块的所有内容。
修正后的代码
import re import os import json def extract_data(text): # 匹配完整的问题区块:编号+标题(Node ID: ...) + 后续的所有-开头的列表项 block_pattern = r'\d+\.\s\*(.*?)\((Node ID:.*?)\)\s*(.*?)(?=\d+\.\s*\*|$)' block_regex = re.compile(block_pattern, re.DOTALL) data = [] blocks = block_regex.findall(text) for block in blocks: heading = block[0].strip() # 提取所有Node ID格式的内容 node_ids = re.findall(r'\d+:\d+', block[1]) # 提取区块内所有以-开头的详情项 detailed_items = re.findall(r'-\s*(.*?)(?=\n-\s|$)', block[2], re.DOTALL) # 清理每个详情项的多余换行和空格 detailed_desc = [item.strip().replace('\n', ' ') for item in detailed_items] data.append({ 'issue': heading, 'node_ids': node_ids, 'detailed_feedback': detailed_desc }) return data def main(): txt_file_path = "./FormatOutput/sample.txt" if os.path.exists(txt_file_path): try: with open(txt_file_path, 'r') as txt_file: text_data = txt_file.read() except Exception as e: print("读取文本文件出错:", e) return else: print("文件不存在:", txt_file_path) return structured_data = extract_data(text_data) json_data = json.dumps(structured_data, indent=4, ensure_ascii=False) print(json_data) if __name__ == "__main__": main()
关键修改说明
- 调整正则逻辑:先匹配完整的问题区块,确保标题、Node ID、详情项属于同一组,避免字段错位。
- 直接提取详情列表:用正则捕获所有以
-开头的列表项,替代sent_tokenize,保证每个列表项作为独立元素。 - 清理文本格式:去除详情项内的多余换行和空格,保证输出文本整洁。
- 优化编码:添加
ensure_ascii=False,支持中文文本正常输出。
预期输出示例
[ { "issue": "内容清晰度与结构优化", "node_ids": ["117:55", "117:135"], "detailed_feedback": [ "将文本整合为清晰段落,说明服务的用途、优势、功能及价值主张。", "修订使命宣言,直接点明服务解决的客户痛点。" ] }, { "issue": "行动号召(CTA)优化", "node_ids": ["117:38", "117:89"], "detailed_feedback": [ "修改“立即购买”CTA,加入服务名称或优惠信息(例如:“立即获取[服务名称]”)。", "调整CTA制造紧迫感:“立即开始免费试用” & “立即保护您的系统”。", "合并相似操作的CTA,区分试用与购买选项。", "调整“免费下载”按钮颜色提升可见度(浅色或对比色)。" ] }, { "issue": "标题与引言增强", "node_ids": ["117:27"], "detailed_feedback": [ "增大标题字号与字重,在产品图上更突出。", "使用更大字号或对比色的项目符号高亮功能。" ] } ]
内容的提问来源于stack exchange,提问作者faizan_bhatti
相关产品推荐
相关产品推荐

