spaCy SpanGroup定义、用法及重叠Span问题的技术咨询
解答
一、SpanGroup 是什么?何时使用?
SpanGroup 是 spaCy v3.0+ 引入的Span容器,用于在单个Doc中管理一组或多组Span,核心特性是支持重叠Span,且没有doc.ents的非重叠约束。
- 适用场景:
- 处理重叠标注:比如你的
food(如"tomato soup")和ingredient(如"tomato")重叠场景 - 分类存储多类型Span:同时维护命名实体、自定义短语、情感区间等多组独立的Span集合
- 替代
doc.ents存储非实体类标注:当你不需要实体的唯一性约束时
- 处理重叠标注:比如你的
二、重叠Span报错的原因
你当前代码中调用了doc.set_ents(span_lst),而doc.ents(实体集合)要求每个Token只能属于一个实体Span,重叠的Span会直接触发验证错误——这正是你遇到的问题。
三、用SpanGroup解决重叠问题的具体方案
完全可以用SpanGroup替代doc.ents来存储重叠Span,以下是修改后的代码及关键说明:
修改后的核心代码
def create_docbin(fname: str, basename: str, nlp): """Create a DocBin from a CSV with rows (list(indexes), text, label)""" doc_bin = DocBin() # 假设CSV每行格式为:[索引列表], 文本, 标签(food/ingredient) for spans, text, label in read_data(fname): ms = make_spans(spans) doc = nlp(text) span_lst = [] for start, end in ms: span = doc.char_span(start, end, label=label.upper()) # 标签转为大写符合spaCy惯例 if span is not None: span_lst.append(span) # 将Span添加到对应标签的SpanGroup if label not in doc.spans: doc.spans[label] = [] doc.spans[label].extend(span_lst) # 可选:也可以把所有Span放到同一个总SpanGroup下 # doc.spans["all_annotations"] = doc.spans.get("all_annotations", []) + span_lst doc_bin.add(doc) doc_bin.to_disk(f'corpus/{basename}.spacy') # 同时需要修改read_data函数以读取标签列 def read_data(fname: str): """Read data from CSV with rows (list(indexes), text, label)""" with open(fname, newline='') as csvfile: reader = csv.reader(csvfile) _ = next(reader) for row in reader: lst = ast.literal_eval(row[0]) text = row[1] label = row[2] yield lst, text, label
关键修改点:
- 移除
doc.set_ents()调用:彻底避开实体的非重叠约束 - 按标签分组存储SpanGroup:
doc.spans["food"]和doc.spans["ingredient"]分别存储两类重叠Span,互不干扰 - 保留原始标签信息:确保后续可以精准访问不同类型的Span集合
四、后续使用SpanGroup的注意事项
- 访问Span:通过
doc.spans["food"]即可获取所有food类型的Span,包括与ingredient重叠的部分 - 模型训练配置:如果要基于这些Span训练模型,需在spaCy配置文件中指定
spans_key(比如["food", "ingredient"]或"all_annotations"),让模型读取对应的SpanGroup标注 - 兼容性:确保使用spaCy v3.0及以上版本,SpanGroup在旧版本中不可用
内容的提问来源于stack exchange,提问作者jason
相关产品推荐
相关产品推荐

