You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何有条件处理JSON数据:literal实体转字符串,metaphoric实体保留结构

NER实体格式转换解决方案

需求说明

处理给定的JSON结构数据,按以下规则转换每个句子中的实体:

  • literal类型实体:直接替换为该实体的word字段纯文本,丢弃原实体的entity_group、score、start、end属性
  • metaphoric类型实体:保留原有完整字典结构,但转换为JSON字符串形式

原始输入数据

{
    "journal.pbio.0050304.xml": {
        "sentence": [
            [
                {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299}
            ],
            [
                {"entity_group": "literal", "score": 0.9932352, "word": "RA, Fgfs, and Wnts are all produced at the posterior of the embryo, and might therefore be expected to form posterior-", "start": 0, "end": 118},
                {"entity_group": "metaphoric", "score": 0.874372, "word": "to", "start": 118, "end": 120},
                {"entity_group": "literal", "score": 0.99049604, "word": "-anterior gradients (for Fgf8", "start": 120, "end": 149},
                {"entity_group": "metaphoric", "score": 0.9993481, "word": "this", "start": 150, "end": 154}
            ]
        ]
    },
    "journal.pbio.0050093.xml": {
        "sentence": [
            [
                {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299}
            ]
        ]
    }
}

原尝试代码问题分析

给出的trial2代码存在两个核心问题:

  1. 循环嵌套逻辑错误:在处理单个XML节点的循环内,就遍历resu的所有值,导致重复处理已完成的文件数据
  2. 未实现metaphoric实体的字符串化处理,仅完成了literal实体的替换

修正后代码

import os
import xml.etree.ElementTree as ET
import json

resu = {}
words_input_dir = "./your_input_dir"  # 替换为实际输入目录

for filename in os.listdir(words_input_dir):
    if filename.endswith(".xml"):
        # 解析XML文件
        tree = ET.parse(os.path.join(words_input_dir, filename))
        root = tree.getroot()
        node = root.findall("./body/sec/p")
        sentences_data = []
        
        # 提取并处理每个段落的实体数据
        for x in node:
            if x is not None and x.text:
                coco = x.text
                data = nerpipeline(str(coco))  # 假设nerpipeline返回句子级的实体列表
                sentences_data.extend(data)
        
        # 按规则转换实体
        processed_sentences = []
        for sentence in sentences_data:
            new_sentence = []
            for entity in sentence:
                if entity["entity_group"] == "literal":
                    new_sentence.append(entity["word"])
                elif entity["entity_group"] == "metaphoric":
                    new_sentence.append(json.dumps(entity))
            processed_sentences.append(new_sentence)
        
        resu[filename] = {"sentence": processed_sentences}

# 输出处理后的结果
print(json.dumps(resu, indent=2))

处理后示例输出

{
  "journal.pbio.0050304.xml": {
    "sentence": [
      [
        "The anterior–posterior (A–P) axis "
      ],
      [
        "RA, Fgfs, and Wnts are all produced at the posterior of the embryo, and might therefore be expected to form posterior-",
        "{\"entity_group\": \"metaphoric\", \"score\": 0.874372, \"word\": \"to\", \"start\": 118, \"end\": 120}",
        "-anterior gradients (for Fgf8",
        "{\"entity_group\": \"metaphoric\", \"score\": 0.9993481, \"word\": \"this\", \"start\": 150, \"end\": 154}"
      ]
    ]
  },
  "journal.pbio.0050093.xml": {
    "sentence": [
      [
        "The anterior–posterior (A–P) axis "
      ]
    ]
  }
}

内容的提问来源于stack exchange,提问作者Idkwhatywantmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 07:15:32