如何处理实体识别结果列表生成指定格式文本?解决代码索引报错
问题背景
我有一份实体识别输出结果,是包含字典的列表,内容如下:
[{'entity_group': 'literal', 'score': 0.99999213, 'word': 'DNA', 'start': 0, 'end': 3}, {'entity_group': 'metaphoric', 'score': 0.9768174, 'word': 'loop', 'start': 4, 'end': 8}, {'entity_group': 'literal', 'score': 0.9039155, 'word': 'ing,', 'start': 8, 'end': 12}]
我需要合并所有literal类型实体的文本,将metaphoric类型实体格式化为{metaphoric:实体内容}的形式,最终得到如下结果:
DNA {metaphoric:loop}ing
我尝试了以下Python代码,但出现“string indices must be integers”的错误,无法得到预期结果:
with open(r'MYFILE.txt', 'r') as res: texty = res.read() for group in texty[::-1]: ent = group["entity_group"] if ent != 'literal': text2 = replace_at(ent, group['end'], group['end'], text) print(text2)
错误原因
- 未解析字符串为Python对象:直接读取文件内容到
texty得到的是字符串,不是字典列表。遍历字符串时group是单个字符,无法用字典索引group["entity_group"],触发索引错误。 - 逻辑与语法问题:
replace_at函数未定义、变量text未声明,且循环内缩进错误,整体逻辑未按实体位置顺序处理。
解决方案
先将文件中的字符串解析为Python列表,按实体的start位置排序保证顺序,再逐个处理实体并拼接结果:
import ast # 读取并解析文件中的实体列表 with open(r'MYFILE.txt', 'r') as res: texty = res.read().strip() entities = ast.literal_eval(texty) # 按实体在原文中的起始位置排序 entities.sort(key=lambda x: x['start']) result = [] for ent in entities: if ent['entity_group'] == 'literal': # 去除literal实体末尾多余的逗号 word = ent['word'].rstrip(',') result.append(word) else: # 格式化metaphoric实体 result.append(f'{{metaphoric:{ent["word"]}}}') # 拼接所有片段得到最终文本 final_text = ''.join(result) print(final_text)
代码说明
ast.literal_eval():安全将字符串形式的列表转为Python对象,避免eval()的安全风险。- 按
start排序:确保实体按原文出现顺序处理,避免拼接混乱。 - 处理逗号:针对原数据中
ing,的情况,用rstrip(',')去除末尾逗号,匹配预期输出。 - 拼接结果:用
''.join()将所有处理后的片段合并成最终字符串。
内容的提问来源于stack exchange,提问作者Idkwhatywantmed
相关产品推荐
相关产品推荐

