You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Stanford CoreNLP的Simple API获取多token实体提及?

关于CoreNLP Simple API获取完整实体提及的解答

当然可以通过Simple API获取完整的多token实体提及啦!其实Simple API里藏着现成的方法,能直接拿到组合好的实体,完全不用你自己去手动拼接token。

具体操作步骤很简单:

  • 先创建Document对象并完成标注
  • 遍历文档里的每个Sentence
  • 调用句子的mentions()方法,这个方法会返回CoreMention对象的列表,每个对象就对应一个完整的实体提及(比如"George Washington"这种多token组合)

给你举个直观的代码例子:

import edu.stanford.nlp.simple.*;

public class NERExample {
    public static void main(String[] args) {
        Document doc = new Document("George Washington was the first president of the United States.");
        for (Sentence sent : doc.sentences()) {
            for (CoreMention mention : sent.mentions()) {
                System.out.println("实体文本: " + mention.text());
                System.out.println("实体类型: " + mention.entityType());
            }
        }
    }
}

运行这段代码后,你就能直接拿到"George Washington"和"United States"这类完整实体,以及它们对应的NER类型(PERSON和LOCATION)。

要是你需要更精准的筛选,还能用mentions()的重载版本,指定只获取特定类型的实体,比如只提取人物类:

for (CoreMention mention : sent.mentions(MentionType.PERSON)) {
    // 只处理人物实体的逻辑
}

这样是不是比自己拼接token省心多啦?

内容的提问来源于stack exchange,提问作者demongolem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:42:55