CoreNLP TokenSequencePattern入门匹配失效问题求助
兄弟,我刚上手CoreNLP的时候也踩过一模一样的坑!你的问题核心就是TokenSequencePattern的匹配语法用错了,再加上可能没正确遍历匹配结果,才会出现“无报错但匹配不到”的情况。
1. 先搞懂匹配单个词的正确语法
TokenSequencePattern不是直接写单词就行的,它需要用{属性名:"属性值"}的格式来指定token的匹配规则。比如你要匹配"This",得写成{word:"This"}——这里的word是token的属性名,对应token的文本内容。你之前直接写单词的话,CoreNLP会把它当成无效的表达式,自然找不到匹配内容。
至于你用[]能匹配到两个句子,是因为[]在TokenSequencePattern里代表空的分组匹配,它会匹配每个句子里的“空序列”,所以看起来像是匹配到了整个句子,但这并不是你真正想要的效果。
2. 给你修正后的完整代码示例
我把正确的写法整理好了,你可以直接参考:
import edu.stanford.nlp.ling.CoreAnnotations; import edu.stanford.nlp.pipeline.Annotation; import edu.stanford.nlp.pipeline.StanfordCoreNLP; import edu.stanford.nlp.util.CoreMap; import edu.stanford.nlp.ling.tokensregex.TokenSequencePattern; import edu.stanford.nlp.ling.tokensregex.TokenSequenceMatcher; import java.util.Properties; import java.util.List; public class CoreNLPTokenMatchTest { public static void main(String[] args) { // 初始化pipeline,至少需要tokenize和sssplit来拆分句子和token Properties props = new Properties(); props.put("annotators", "tokenize, ssplit"); StanfordCoreNLP pipeline = new StanfordCoreNLP(props); // 待处理的文本 Annotation document = new Annotation("This is my first test. Another sentence here."); // 让pipeline处理文档 pipeline.annotate(document); // 获取所有拆分后的句子 List<CoreMap> sentences = document.get(CoreAnnotations.SentencesAnnotation.class); // 创建匹配单个词"This"的正确模式 TokenSequencePattern pattern = TokenSequencePattern.compile("{word:\"This\"}"); // 遍历每个句子查找匹配 for (CoreMap sentence : sentences) { TokenSequenceMatcher matcher = pattern.getMatcher(sentence); // 循环查找所有匹配结果(避免只找第一个就停) while (matcher.find()) { System.out.println("匹配到的内容: " + matcher.group()); // 也可以直接获取匹配到的token对象 System.out.println("对应的token: " + matcher.groupNodes().get(0)); } } } }
3. 额外的小技巧
- 如果要忽略大小写匹配,可以用正则模式:
{word:/This/i},这样"This"和"this"都能被匹配到; - 如果要匹配词性(比如名词),需要在annotators里加上
pos,然后用{pos:"NN"}这样的表达式。
换成正确的匹配格式后,你就能正常匹配到目标单词了!
内容的提问来源于stack exchange,提问作者webber
相关产品推荐
相关产品推荐

