You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:修复LLM项目中Python JSON输出格式化问题

问题:LLM反馈文本转JSON格式异常修复

问题描述

开发LLM项目时,需要将LLM生成的反馈响应格式化为指定结构的JSON。现有Python代码通过正则表达式提取标题(issue)、node_ids、详细描述(detailed_feedback),但输出格式不符合预期,无法正确关联每个问题的对应字段。

原代码

import re
import os
import json
from nltk.tokenize import sent_tokenize

def extract_data(text):
    section_pattern = r'\d+\.\s\*(.*?)\((Node ID:.*?)\).*?((?=\d+\.\s*\*)|$)|\-\s(.*?)\n\n'
    section_regex = re.compile(section_pattern, re.MULTILINE | re.DOTALL)
    
    matches = section_regex.findall(text)
    data = []
    # print(matches)
    
    for match in matches:
        heading = match[0].strip()
        node_ids = re.findall(r'\d+:\d+', match[1])
        detailed_desc = extract_detailed_desc(match[3].strip())
        data.append({
                'issue': heading,
                'node_ids': node_ids,
                'detailed_feedback': detailed_desc
                })

    return data

def extract_detailed_desc(text):
    
    sentences = sent_tokenize(text)
    detailed_desc = []
    for sentence in sentences:
        detailed_desc.append(sentence.strip('-').strip())
    return detailed_desc


def main():
    txt_file_path = "./FormatOutput/sample.txt"

    if os.path.exists(txt_file_path):
        try:
            with open(txt_file_path, 'r') as txt_file:
                data = txt_file.read()
                # print(data)
        except Exception as e:
            print("Error occurred while reading the text file:", e)
    else:
        print("File not found:", txt_file_path)

    structured_data = extract_data(data)
    json_data = json.dumps(structured_data, indent=4)
    print(json_data)

if __name__ == "__main__":
    main()

问题分析

  1. 正则表达式采用分支匹配,将问题标题和详情项分开捕获,导致标题与详情无法正确关联,出现字段缺失或错位。
  2. 使用sent_tokenize拆分详情文本,会把单个列表项拆分成多个句子,不符合预期的详情列表结构。
  3. 正则的边界匹配逻辑不严谨,无法完整捕获每个问题区块的所有内容。

修正后的代码

import re
import os
import json

def extract_data(text):
    # 匹配完整的问题区块:编号+标题(Node ID: ...) + 后续的所有-开头的列表项
    block_pattern = r'\d+\.\s\*(.*?)\((Node ID:.*?)\)\s*(.*?)(?=\d+\.\s*\*|$)'
    block_regex = re.compile(block_pattern, re.DOTALL)
    
    data = []
    blocks = block_regex.findall(text)
    
    for block in blocks:
        heading = block[0].strip()
        # 提取所有Node ID格式的内容
        node_ids = re.findall(r'\d+:\d+', block[1])
        # 提取区块内所有以-开头的详情项
        detailed_items = re.findall(r'-\s*(.*?)(?=\n-\s|$)', block[2], re.DOTALL)
        # 清理每个详情项的多余换行和空格
        detailed_desc = [item.strip().replace('\n', ' ') for item in detailed_items]
        
        data.append({
            'issue': heading,
            'node_ids': node_ids,
            'detailed_feedback': detailed_desc
        })

    return data

def main():
    txt_file_path = "./FormatOutput/sample.txt"

    if os.path.exists(txt_file_path):
        try:
            with open(txt_file_path, 'r') as txt_file:
                text_data = txt_file.read()
        except Exception as e:
            print("读取文本文件出错:", e)
            return
    else:
        print("文件不存在:", txt_file_path)
        return

    structured_data = extract_data(text_data)
    json_data = json.dumps(structured_data, indent=4, ensure_ascii=False)
    print(json_data)

if __name__ == "__main__":
    main()

关键修改说明

  1. 调整正则逻辑:先匹配完整的问题区块,确保标题、Node ID、详情项属于同一组,避免字段错位。
  2. 直接提取详情列表:用正则捕获所有以- 开头的列表项,替代sent_tokenize,保证每个列表项作为独立元素。
  3. 清理文本格式:去除详情项内的多余换行和空格,保证输出文本整洁。
  4. 优化编码:添加ensure_ascii=False,支持中文文本正常输出。

预期输出示例

[
    {
        "issue": "内容清晰度与结构优化",
        "node_ids": ["117:55", "117:135"],
        "detailed_feedback": [
            "将文本整合为清晰段落,说明服务的用途、优势、功能及价值主张。",
            "修订使命宣言,直接点明服务解决的客户痛点。"
        ]
    },
    {
        "issue": "行动号召(CTA)优化",
        "node_ids": ["117:38", "117:89"],
        "detailed_feedback": [
            "修改“立即购买”CTA,加入服务名称或优惠信息(例如:“立即获取[服务名称]”)。",
            "调整CTA制造紧迫感:“立即开始免费试用” & “立即保护您的系统”。",
            "合并相似操作的CTA,区分试用与购买选项。",
            "调整“免费下载”按钮颜色提升可见度(浅色或对比色)。"
        ]
    },
    {
        "issue": "标题与引言增强",
        "node_ids": ["117:27"],
        "detailed_feedback": [
            "增大标题字号与字重,在产品图上更突出。",
            "使用更大字号或对比色的项目符号高亮功能。"
        ]
    }
]

内容的提问来源于stack exchange,提问作者faizan_bhatti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 10:46:36