基于Stanford Stanza NLP:获取指定实体跨度的所有词ID
用Stanza匹配命名实体对应的分词
我正在使用Stanza的处理管道提取词与命名实体:
sentence.entities返回包含起始、结束字符位置的命名实体列表,示例如下:
{ "text": "Dante Alighieri", "type": "PER", "start_char": 1, "end_char": 16 }
sentence.words返回包含起始、结束字符位置的全部分词列表,对应示例片段如下:
{ "id": 1, "text": "Dante", "lemma": "Dante", "upos": "PROPN", "xpos": "SP", "head": 3, "deprel": "nsubj", "start_char": 1, "end_char": 6 } { "id": 2, "text": "Alighieri", "lemma": "Alighieri", "upos": "PROPN", "xpos": "SP", "head": 1, "deprel": "flat:name", "start_char": 7, "end_char": 16 } { "id": 3, "text": "scrisse", "lemma": "scrivere", "upos": "VERB", "xpos": "V", "feats": "Mood=Ind|Number=Sing|Person=3|Tense=Past|VerbForm=Fin", "head": 0, "deprel": "root", "start_char": 17, "end_char": 24 }
我需要生成包含在命名实体字符跨度内的所有词的列表,以上述示例为例,应返回ID为1和2的词,而非ID为3的词。
内容的提问来源于stack exchange,提问作者Robert Alexander
相关产品推荐
相关产品推荐

