Spacy 2.2.4中Displacy manual=True模式报错KeyError: 'ents'怎么解决?
使用spaCy displacy手动渲染实体时出现KeyError: 'ents'
在spaCy 2.2.4版本中,执行以下代码时触发KeyError: 'ents':
from spacy import displacy ex = { "text": "Welcome to the Bank of China.", "spans": [ {"start_token": 3, "end_token": 6, "label": "ORG"}, {"start_token": 5, "end_token": 6, "label": "GPE"}, ], "tokens": ["Welcome", "to", "the", "Bank", "of", "China", "."], } html = displacy.render(ex, style="ent", manual=True)
报错栈信息:
KeyError Traceback (most recent call last) Input In [20], in <cell line: 13>() 3 from spacy import displacy 5 ex = { 6 "text": "Welcome to the Bank of China.", 7 "spans": [ (...) 11 "tokens": ["Welcome", "to", "the", "Bank", "of", "China", "."], 12 } ---> 13 html = displacy.render(ex, style="ent", manual=True) File /data/venv/lib/python3.8/site-packages/spacy/displacy/__init__.py:52, in render(docs, style, page, minify, jupyter, options, manual) 50 renderer = renderer(options=options) 51 parsed = [converter(doc, options) for doc in docs] if not manual else docs ---> 52 _html["parsed"] = renderer.render(parsed, page=page, minify=minify).strip() 53 html = _html["parsed"] 54 if RENDER_WRAPPER is not None: File /data/venv/lib/python3.8/site-packages/spacy/displacy/render.py:280, in EntityRenderer.render(self, parsed, page, minify) 278 self.direction = settings.get("direction", DEFAULT_DIR) 279 self.lang = settings.get("lang", DEFAULT_LANG) ---> 280 rendered.append(self.render_ents(p["text"], p["ents"], p.get("title"))) 281 if page: 282 docs = "".join([TPL_FIGURE.format(content=doc) for doc in rendered]) KeyError: 'ents'
问题原因
你参考的是spaCy v3+版本的官方文档,但当前使用的是v2.2.4旧版本。v2.x的displacy手动渲染实体时,不支持spans、start_token、end_token这些v3+新增的字段,而是要求使用ents字段,并且实体范围必须用字符的起始/结束位置(而非token索引)来定义。
修复方案
方案1:手动指定字符位置
直接替换字段结构,将spans改为ents,并填写对应实体的字符起止位置:
from spacy import displacy ex = { "text": "Welcome to the Bank of China.", "ents": [ {"start": 12, "end": 23, "label": "ORG"}, # 对应"Bank of China"的字符范围 {"start": 19, "end": 23, "label": "GPE"}, # 对应"China"的字符范围 ] } html = displacy.render(ex, style="ent", manual=True)
方案2:通过token自动计算字符位置
如果不想手动计算字符位置,可以基于token列表自动推导:
from spacy import displacy text = "Welcome to the Bank of China." tokens = ["Welcome", "to", "the", "Bank", "of", "China", "."] # 计算每个token的字符起始位置 token_starts = [] current_pos = 0 for token in tokens: token_starts.append(current_pos) current_pos += len(token) + 1 # 加1是因为token之间存在空格 ex = { "text": text, "ents": [ # start_token=3对应token_starts[3],end_token=6对应token_starts[6](end为排他索引) {"start": token_starts[3], "end": token_starts[6], "label": "ORG"}, {"start": token_starts[5], "end": token_starts[6], "label": "GPE"}, ] } html = displacy.render(ex, style="ent", manual=True)
内容的提问来源于stack exchange,提问作者Kaushik J
相关产品推荐
相关产品推荐

