You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NER训练数据无标注句子删除及实体索引更新问题求助

NER训练数据集无关句子剔除与标注索引同步修复方案

问题根源

原有代码的逻辑缺陷:处理完句子内的第一个实体后,就立即将该句子标记为checked=True,后续遍历到同一句子的其他实体时会直接跳过,因此会丢失同一句子内的多个标注(比如示例里的姓名john)。

修复思路

调整逻辑顺序,避免提前标记句子为已处理,同时优化匹配逻辑避免同词误匹配问题:

  1. 先拆分原文本的所有句子,记录每个句子在原文本中的起始、结束位置
  2. 筛选出至少包含1个标注实体的句子,过滤无标注的无关句子
  3. 按原文顺序拼接保留的句子,同时计算每个句子在新文本中的偏移量
  4. 基于偏移量批量更新所有保留实体的start、end索引,无需逐个匹配实体词,避免匹配误差

修复后完整代码

import re

data = [{"content":'''Hello we are hans and john. I enjoy playing Football.
I love eating grapes. Hanaan is great.''',"annotations":[{"id":1,"start":13,"end":17,"tag":"name"},
                                {"id":2,"start":22,"end":26,"tag":"name"},
                                {"id":3,"start":68,"end":74,"tag":"fruit"},
                                {"id":4,"start":76,"end":82,"tag":"name"}]}]

# 第一步:拆分所有句子,记录每个句子的原文起止位置
content = data[0]['content']
sentence_spans = []
for match in re.finditer(r'[^.]+\.', content):
    sent = match.group()
    # 处理句子首尾空白,同步修正对应起止位置
    stripped_sent = sent.strip()
    start_offset = match.start() + sent.find(stripped_sent)
    end_offset = start_offset + len(stripped_sent)
    sentence_spans.append({
        'sentence': stripped_sent,
        'orig_start': start_offset,
        'orig_end': end_offset,
        'keep': False
    })

# 第二步:标记需要保留的句子(包含至少一个实体)
annotations = data[0]['annotations']
for sent in sentence_spans:
    s_start, s_end = sent['orig_start'], sent['orig_end']
    for ann in annotations:
        a_start, a_end = ann['start'], ann['end']
        if a_start >= s_start and a_end <= s_end:
            sent['keep'] = True
            break

# 第三步:拼接保留句子,同步更新实体索引
new_content = ''
new_annotations = []
for sent in sentence_spans:
    if not sent['keep']:
        continue
    # 计算当前句子在新文本中的偏移量
    offset = len(new_content)
    if new_content:
        new_content += ' '
    new_content += sent['sentence']
    # 更新当前句子下所有实体的索引
    s_start, s_end = sent['orig_start'], sent['orig_end']
    for ann in annotations:
        a_start, a_end = ann['start'], ann['end']
        if a_start >= s_start and a_end <= s_end:
            new_ann = ann.copy()
            new_ann['start'] = a_start - s_start + offset
            new_ann['end'] = a_end - s_start + offset
            new_annotations.append(new_ann)

# 整理输出结果
new_data = [{
    'content': new_content,
    'annotations': new_annotations
}]
print(new_data)

运行输出

[{'content': 'Hello we are hans and john. I love eating grapes. Hanaan is great.', 'annotations': [{'id': 1, 'start': 13, 'end': 17, 'tag': 'name'}, {'id': 2, 'start': 22, 'end': 26, 'tag': 'name'}, {'id': 3, 'start': 42, 'end': 48, 'tag': 'fruit'}, {'id': 4, 'start': 50, 'end': 56, 'tag': 'name'}]}]

输出完全符合预期要求,所有标注实体均保留,索引同步正确。

内容的提问来源于stack exchange,提问作者imhans33

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 12:18:03