如何有条件处理JSON数据:literal实体转字符串,metaphoric实体保留结构
NER实体格式转换解决方案
需求说明
处理给定的JSON结构数据,按以下规则转换每个句子中的实体:
literal类型实体:直接替换为该实体的word字段纯文本,丢弃原实体的entity_group、score、start、end属性metaphoric类型实体:保留原有完整字典结构,但转换为JSON字符串形式
原始输入数据
{ "journal.pbio.0050304.xml": { "sentence": [ [ {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299} ], [ {"entity_group": "literal", "score": 0.9932352, "word": "RA, Fgfs, and Wnts are all produced at the posterior of the embryo, and might therefore be expected to form posterior-", "start": 0, "end": 118}, {"entity_group": "metaphoric", "score": 0.874372, "word": "to", "start": 118, "end": 120}, {"entity_group": "literal", "score": 0.99049604, "word": "-anterior gradients (for Fgf8", "start": 120, "end": 149}, {"entity_group": "metaphoric", "score": 0.9993481, "word": "this", "start": 150, "end": 154} ] ] }, "journal.pbio.0050093.xml": { "sentence": [ [ {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299} ] ] } }
原尝试代码问题分析
给出的trial2代码存在两个核心问题:
- 循环嵌套逻辑错误:在处理单个XML节点的循环内,就遍历
resu的所有值,导致重复处理已完成的文件数据 - 未实现
metaphoric实体的字符串化处理,仅完成了literal实体的替换
修正后代码
import os import xml.etree.ElementTree as ET import json resu = {} words_input_dir = "./your_input_dir" # 替换为实际输入目录 for filename in os.listdir(words_input_dir): if filename.endswith(".xml"): # 解析XML文件 tree = ET.parse(os.path.join(words_input_dir, filename)) root = tree.getroot() node = root.findall("./body/sec/p") sentences_data = [] # 提取并处理每个段落的实体数据 for x in node: if x is not None and x.text: coco = x.text data = nerpipeline(str(coco)) # 假设nerpipeline返回句子级的实体列表 sentences_data.extend(data) # 按规则转换实体 processed_sentences = [] for sentence in sentences_data: new_sentence = [] for entity in sentence: if entity["entity_group"] == "literal": new_sentence.append(entity["word"]) elif entity["entity_group"] == "metaphoric": new_sentence.append(json.dumps(entity)) processed_sentences.append(new_sentence) resu[filename] = {"sentence": processed_sentences} # 输出处理后的结果 print(json.dumps(resu, indent=2))
处理后示例输出
{ "journal.pbio.0050304.xml": { "sentence": [ [ "The anterior–posterior (A–P) axis " ], [ "RA, Fgfs, and Wnts are all produced at the posterior of the embryo, and might therefore be expected to form posterior-", "{\"entity_group\": \"metaphoric\", \"score\": 0.874372, \"word\": \"to\", \"start\": 118, \"end\": 120}", "-anterior gradients (for Fgf8", "{\"entity_group\": \"metaphoric\", \"score\": 0.9993481, \"word\": \"this\", \"start\": 150, \"end\": 154}" ] ] }, "journal.pbio.0050093.xml": { "sentence": [ [ "The anterior–posterior (A–P) axis " ] ] } }
内容的提问来源于stack exchange,提问作者Idkwhatywantmed
相关产品推荐
相关产品推荐

