spaCy中同一文本标注多实体类型时如何删除指定实体及跨度?
spaCy同跨度多实体标注删除方法
问题说明
在spaCy的实体标注数据处理中,若同一个文本跨度被标注了多种不同实体类型,可通过过滤标注列表的方式删除指定实体类型的对应标注项。
输入示例
[('PERSON1: thank you for calling ABCD my name is QWERT and him i speaking with PERSON2: hi QWERT this is QAZWSXE RFVTGB', {'entities': [(102, 116, 'PERSON'), (29, 33, 'ORG'), (45, 50, 'PERSON'), (45, 50, 'ORG')]})]
上述示例中,跨度(45, 50)对应的文本片段QWERT被同时标注为PERSON和ORG两种实体类型,需要删除(45, 50, 'ORG')标注项。
实现代码
针对指定标注项删除
如果明确知道要删除的跨度和实体类型,可直接用列表推导式过滤:
# 原始标注数据 raw_data = [('PERSON1: thank you for calling ABCD my name is QWERT and him i speaking with PERSON2: hi QWERT this is QAZWSXE RFVTGB', {'entities': [(102, 116, 'PERSON'), (29, 33, 'ORG'), (45, 50, 'PERSON'), (45, 50, 'ORG')]})] # 定义待删除的跨度和对应实体类型 to_remove = (45, 50, 'ORG') # 过滤实体列表 for item in raw_data: item[1]['entities'] = [ent for ent in item[1]['entities'] if ent != to_remove] print(raw_data)
通用批量去重(同跨度保留高优先级实体)
如果需要批量处理所有同跨度多标注的情况,可以提前定义实体优先级,同一个跨度只保留优先级最高的实体:
# 实体优先级,数值越小优先级越高 ent_priority = { 'PERSON': 1, 'ORG': 2, # 可补充其他实体类型的优先级 } raw_data = [('PERSON1: thank you for calling ABCD my name is QWERT and him i speaking with PERSON2: hi QWERT this is QAZWSXE RFVTGB', {'entities': [(102, 116, 'PERSON'), (29, 33, 'ORG'), (45, 50, 'PERSON'), (45, 50, 'ORG')]})] for item in raw_data: # 按跨度分组,存储每个跨度下优先级最高的实体 span_map = {} for ent in item[1]['entities']: start, end, label = ent span_key = (start, end) # 如果该跨度还没存,或者当前实体优先级更高,就替换 if span_key not in span_map or ent_priority[label] < ent_priority[span_map[span_key][2]]: span_map[span_key] = ent # 替换原实体列表 item[1]['entities'] = list(span_map.values()) print(raw_data)
输出结果
两种方法运行后都可得到期望的输出:
[('PERSON1: thank you for calling ABCD my name is QWERT and him i speaking with PERSON2: hi QWERT this is QAZWSXE RFVTGB', {'entities': [(102, 116, 'PERSON'), (29, 33, 'ORG'), (45, 50, 'PERSON')]})]
内容的提问来源于stack exchange,提问作者Jatin Chaturvedi
相关产品推荐
相关产品推荐

