如何通过Stanford CoreNLP的Simple API获取多token实体提及?
关于CoreNLP Simple API获取完整实体提及的解答
当然可以通过Simple API获取完整的多token实体提及啦!其实Simple API里藏着现成的方法,能直接拿到组合好的实体,完全不用你自己去手动拼接token。
具体操作步骤很简单:
- 先创建
Document对象并完成标注 - 遍历文档里的每个
Sentence - 调用句子的
mentions()方法,这个方法会返回CoreMention对象的列表,每个对象就对应一个完整的实体提及(比如"George Washington"这种多token组合)
给你举个直观的代码例子:
import edu.stanford.nlp.simple.*; public class NERExample { public static void main(String[] args) { Document doc = new Document("George Washington was the first president of the United States."); for (Sentence sent : doc.sentences()) { for (CoreMention mention : sent.mentions()) { System.out.println("实体文本: " + mention.text()); System.out.println("实体类型: " + mention.entityType()); } } } }
运行这段代码后,你就能直接拿到"George Washington"和"United States"这类完整实体,以及它们对应的NER类型(PERSON和LOCATION)。
要是你需要更精准的筛选,还能用mentions()的重载版本,指定只获取特定类型的实体,比如只提取人物类:
for (CoreMention mention : sent.mentions(MentionType.PERSON)) { // 只处理人物实体的逻辑 }
这样是不是比自己拼接token省心多啦?
内容的提问来源于stack exchange,提问作者demongolem
相关产品推荐
相关产品推荐

