You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RegexNERAnnotator无法识别撇号问题求助(Stanford JavaNLP)

Fixing Stanford RegexNERAnnotator's Apostrophe Matching Issue

Got it, let's tackle why your Bachelor's of Arts isn't being recognized as a DEGREE entity—this is a common gotcha with token-based NER tools like Stanford CoreNLP!

The Root Cause

Stanford CoreNLP's tokenizer splits contractions like Bachelor's into two separate tokens: Bachelor and 's. Your original regex rule Bachelor's of (Arts|Laws|Science|Engineering) tries to match a single token sequence that doesn't exist after tokenization. When you used Bachelor of Arts, that's three distinct tokens (Bachelor, of, Arts) which aligns perfectly with your initial rule.

Step-by-Step Solution

  1. Adjust your regex mapping file to match the tokenized structure
    Update your regexfile.txt rule to account for the split tokens. Since Bachelor's becomes two tokens, list them separately (RegexNER uses spaces to denote token boundaries):

    Bachelor 's of (Arts|Laws|Science|Engineering) DEGREE Lalor LOCATION PERSON Labor ORGANIZATION
    

    This tells RegexNER to look for the sequence: Bachelor → 's → of → [one of the degree subjects], which matches exactly how the tokenizer processes Bachelor's of Arts.

  2. Verify the tokenization (optional but helpful)
    To confirm how your text is split into tokens, add a quick debug snippet to your code:

    Annotation doc = new Annotation("Bachelor's of Arts");
    pipeline.annotate(doc);
    // Print each token's text
    for (CoreLabel token : doc.get(CoreAnnotations.TokensAnnotation.class)) {
        String tokenText = token.get(CoreAnnotations.TextAnnotation.class);
        System.out.println("Token: " + tokenText);
    }
    

    Running this will output:

    Token: Bachelor
    Token: 's
    Token: of
    Token: Arts
    

    This confirms the exact token sequence you need to target in your regex rule.

  3. Bonus: Match both variants with one rule
    If you want to cover both Bachelor of Arts and Bachelor's of Arts in a single rule, make the 's token optional using parentheses and a ? quantifier:

    Bachelor ('s)? of (Arts|Laws|Science|Engineering) DEGREE Lalor LOCATION PERSON Labor ORGANIZATION
    

    The ('s)? tells RegexNER that the 's token is optional, so it will match both phrasing styles.

Why This Works

RegexNER operates on token sequences, not raw character strings. Every space in your regex rule corresponds to a token boundary in the processed text. By aligning your rule with the tokenizer's output, you ensure the pattern matches exactly what the pipeline sees after tokenization.

内容的提问来源于stack exchange,提问作者Gavy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:31:22