如何将指定JSON转换为含entity_group、start、end的Python元组
提取JSON实体信息并转换为Python元组的实现方法
需求说明
给定如下JSON数据:
{ "journal.pbio.0050304.xml": { "sentence": [ [ {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299} ], [ {"entity_group": "literal", "score": 0.9932352, "word": "RA, Fgfs, and Wnts are all produced at the posterior of the embryo, and might therefore be expected to form posterior-", "start": 0, "end": 118}, {"entity_group": "metaphoric", "score": 0.874372, "word": "to", "start": 118, "end": 120}, {"entity_group": "literal", "score": 0.99049604, "word": "-anterior gradients (for Fgf8", "start": 120, "end": 149}, {"entity_group": "metaphoric", "score": 0.9993481, "word": "this", "start": 150, "end": 154} ] ] }, "journal.pbio.0050093.xml": { "sentence": [ [ {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299} ] ] } }
需要提取每个实体的entity_group、start和end字段,转换为类似[(0, 299, 'literal'), (118, 120, 'metaphoric')]格式的Python元组列表。
实现代码
import json # 加载JSON数据(如果是从文件读取,可替换为with open(...) as f: data = json.load(f)) json_str = ''' { "journal.pbio.0050304.xml": { "sentence": [ [ {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299} ], [ {"entity_group": "literal", "score": 0.9932352, "word": "RA, Fgfs, and Wnts are all produced at the posterior of the embryo, and might therefore be expected to form posterior-", "start": 0, "end": 118}, {"entity_group": "metaphoric", "score": 0.874372, "word": "to", "start": 118, "end": 120}, {"entity_group": "literal", "score": 0.99049604, "word": "-anterior gradients (for Fgf8", "start": 120, "end": 149}, {"entity_group": "metaphoric", "score": 0.9993481, "word": "this", "start": 150, "end": 154} ] ] }, "journal.pbio.0050093.xml": { "sentence": [ [ {"entity_group": "literal", "score": 0.9961686, "word": "The anterior–posterior (A–P) axis ", "start": 0, "end": 299} ] ] } } ''' # 解析JSON数据 data = json.loads(json_str) # 初始化结果列表 result_tuples = [] # 遍历所有文件的实体信息 for content in data.values(): # 遍历每个句子的实体集合 for sentence_entities in content["sentence"]: # 遍历单个实体,提取所需字段生成元组 for entity in sentence_entities: # 按(start, end, entity_group)顺序生成元组,可根据需求调整顺序 result_tuples.append((entity["start"], entity["end"], entity["entity_group"])) # 输出结果 print(result_tuples)
运行结果
[(0, 299, 'literal'), (0, 118, 'literal'), (118, 120, 'metaphoric'), (120, 149, 'literal'), (150, 154, 'metaphoric'), (0, 299, 'literal')]
如果需要将entity_group放在元组首位,只需修改生成元组的代码为:
result_tuples.append((entity["entity_group"], entity["start"], entity["end"]))
对应的输出会变成:
[('literal', 0, 299), ('literal', 0, 118), ('metaphoric', 118, 120), ('literal', 120, 149), ('metaphoric', 150, 154), ('literal', 0, 299)]
内容的提问来源于stack exchange,提问作者Idkwhatywantmed
相关产品推荐
相关产品推荐

