RegexNERAnnotator无法识别撇号问题求助(Stanford JavaNLP)
Got it, let's tackle why your Bachelor's of Arts isn't being recognized as a DEGREE entity—this is a common gotcha with token-based NER tools like Stanford CoreNLP!
The Root Cause
Stanford CoreNLP's tokenizer splits contractions like Bachelor's into two separate tokens: Bachelor and 's. Your original regex rule Bachelor's of (Arts|Laws|Science|Engineering) tries to match a single token sequence that doesn't exist after tokenization. When you used Bachelor of Arts, that's three distinct tokens (Bachelor, of, Arts) which aligns perfectly with your initial rule.
Step-by-Step Solution
Adjust your regex mapping file to match the tokenized structure
Update yourregexfile.txtrule to account for the split tokens. SinceBachelor'sbecomes two tokens, list them separately (RegexNER uses spaces to denote token boundaries):Bachelor 's of (Arts|Laws|Science|Engineering) DEGREE Lalor LOCATION PERSON Labor ORGANIZATIONThis tells RegexNER to look for the sequence:
Bachelor→'s→of→ [one of the degree subjects], which matches exactly how the tokenizer processesBachelor's of Arts.Verify the tokenization (optional but helpful)
To confirm how your text is split into tokens, add a quick debug snippet to your code:Annotation doc = new Annotation("Bachelor's of Arts"); pipeline.annotate(doc); // Print each token's text for (CoreLabel token : doc.get(CoreAnnotations.TokensAnnotation.class)) { String tokenText = token.get(CoreAnnotations.TextAnnotation.class); System.out.println("Token: " + tokenText); }Running this will output:
Token: Bachelor Token: 's Token: of Token: ArtsThis confirms the exact token sequence you need to target in your regex rule.
Bonus: Match both variants with one rule
If you want to cover bothBachelor of ArtsandBachelor's of Artsin a single rule, make the'stoken optional using parentheses and a?quantifier:Bachelor ('s)? of (Arts|Laws|Science|Engineering) DEGREE Lalor LOCATION PERSON Labor ORGANIZATIONThe
('s)?tells RegexNER that the'stoken is optional, so it will match both phrasing styles.
Why This Works
RegexNER operates on token sequences, not raw character strings. Every space in your regex rule corresponds to a token boundary in the processed text. By aligning your rule with the tokenizer's output, you ensure the pattern matches exactly what the pipeline sees after tokenization.
内容的提问来源于stack exchange,提问作者Gavy

