You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Stanford Stanza NLP:获取指定实体跨度的所有词ID

用Stanza匹配命名实体对应的分词

我正在使用Stanza的处理管道提取词与命名实体:

  • sentence.entities返回包含起始、结束字符位置的命名实体列表,示例如下:
{
  "text": "Dante Alighieri",
  "type": "PER",
  "start_char": 1,
  "end_char": 16
}
  • sentence.words返回包含起始、结束字符位置的全部分词列表,对应示例片段如下:
{
  "id": 1,
  "text": "Dante",
  "lemma": "Dante",
  "upos": "PROPN",
  "xpos": "SP",
  "head": 3,
  "deprel": "nsubj",
  "start_char": 1,
  "end_char": 6
}
{
  "id": 2,
  "text": "Alighieri",
  "lemma": "Alighieri",
  "upos": "PROPN",
  "xpos": "SP",
  "head": 1,
  "deprel": "flat:name",
  "start_char": 7,
  "end_char": 16
}
{
  "id": 3,
  "text": "scrisse",
  "lemma": "scrivere",
  "upos": "VERB",
  "xpos": "V",
  "feats": "Mood=Ind|Number=Sing|Person=3|Tense=Past|VerbForm=Fin",
  "head": 0,
  "deprel": "root",
  "start_char": 17,
  "end_char": 24
}

我需要生成包含在命名实体字符跨度内的所有词的列表,以上述示例为例,应返回ID为1和2的词,而非ID为3的词。


内容的提问来源于stack exchange,提问作者Robert Alexander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 11:35:24