You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何处理实体识别结果列表生成指定格式文本?解决代码索引报错

问题背景

我有一份实体识别输出结果,是包含字典的列表,内容如下:

[{'entity_group': 'literal', 'score': 0.99999213, 'word': 'DNA', 'start': 0, 'end': 3}, {'entity_group': 'metaphoric', 'score': 0.9768174, 'word': 'loop', 'start': 4, 'end': 8}, {'entity_group': 'literal', 'score': 0.9039155, 'word': 'ing,', 'start': 8, 'end': 12}]

我需要合并所有literal类型实体的文本,将metaphoric类型实体格式化为{metaphoric:实体内容}的形式,最终得到如下结果:

DNA {metaphoric:loop}ing

我尝试了以下Python代码,但出现“string indices must be integers”的错误,无法得到预期结果:

with open(r'MYFILE.txt', 'r') as res:
  texty = res.read()
  for group in texty[::-1]:
      ent = group["entity_group"]
      if ent != 'literal': 
      text2 = replace_at(ent, group['end'], group['end'], text)
print(text2)

错误原因

  1. 未解析字符串为Python对象:直接读取文件内容到texty得到的是字符串,不是字典列表。遍历字符串时group是单个字符,无法用字典索引group["entity_group"],触发索引错误。
  2. 逻辑与语法问题:replace_at函数未定义、变量text未声明,且循环内缩进错误,整体逻辑未按实体位置顺序处理。

解决方案

先将文件中的字符串解析为Python列表,按实体的start位置排序保证顺序,再逐个处理实体并拼接结果:

import ast

# 读取并解析文件中的实体列表
with open(r'MYFILE.txt', 'r') as res:
    texty = res.read().strip()
    entities = ast.literal_eval(texty)

# 按实体在原文中的起始位置排序
entities.sort(key=lambda x: x['start'])

result = []
for ent in entities:
    if ent['entity_group'] == 'literal':
        # 去除literal实体末尾多余的逗号
        word = ent['word'].rstrip(',')
        result.append(word)
    else:
        # 格式化metaphoric实体
        result.append(f'{{metaphoric:{ent["word"]}}}')

# 拼接所有片段得到最终文本
final_text = ''.join(result)
print(final_text)

代码说明

  • ast.literal_eval():安全将字符串形式的列表转为Python对象,避免eval()的安全风险。
  • 按start排序:确保实体按原文出现顺序处理,避免拼接混乱。
  • 处理逗号:针对原数据中ing,的情况,用rstrip(',')去除末尾逗号,匹配预期输出。
  • 拼接结果:用''.join()将所有处理后的片段合并成最终字符串。

内容的提问来源于stack exchange,提问作者Idkwhatywantmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:45:39