为Spacy Tokenizer添加n't例外:n't转not失败的问题求解
Hey there! I get it—tweaking SpaCy's tokenizer to handle contractions exactly how you want can feel tricky. Let’s break down why your current approach might not be working and walk through two solid solutions to fix it.
First, a quick recap of SpaCy’s default behavior: The English tokenizer splits contractions like "don’t" into two separate tokens: do and n’t. If your attempts to replace "n’t" with "not" aren’t sticking, it’s likely because you’re either missing a step in the pipeline or not overriding SpaCy’s built-in contraction rules correctly.
Solution 1: Add a Custom Pipeline Component (Easiest & Most Flexible)
This approach modifies tokens right after they’re split, which is straightforward and doesn’t require messing with SpaCy’s core tokenizer logic. Here’s how to implement it:
import spacy nlp = spacy.load("en_core_web_sm") # Define a component to swap "n't" tokens with "not" def replace_nt_with_not(doc): for token in doc: if token.text == "n't": token.text = "not" token.lemma_ = "not" # Optional: Update lemma for consistency return doc # Insert the component early in the pipeline (before parser/tagger) nlp.add_pipe(replace_nt_with_not, before="parser") # Test it out doc = nlp("I don't like coffee. She isn't coming.") print([token.text for token in doc]) # Output: ["I", "do", "not", "like", "coffee", ".", "She", "is", "not", "coming", "."]
This works because the component runs immediately after tokenization, so every "n’t" token gets replaced before any downstream processing (like tagging or parsing) happens.
Solution 2: Modify Tokenizer Exceptions (Change Split Behavior)
If you want SpaCy to never split "n’t" as a separate token and instead tokenize contractions like "don’t" directly into "do" + "not", you can override the tokenizer’s exception rules. This requires explicitly defining each contraction you want to handle:
import spacy from spacy.lang.en import English from spacy.symbols import ORTH # Initialize a base tokenizer nlp = English() tokenizer = nlp.tokenizer # Map contractions to their expanded token splits contraction_mappings = { "don't": ["do", "not"], "isn't": ["is", "not"], "wasn't": ["was", "not"], "haven't": ["have", "not"], # Add all other contractions you need here } # Update the tokenizer's special cases for contraction, split_tokens in contraction_mappings.items(): tokenizer.add_special_case(contraction, [{ORTH: token} for token in split_tokens]) # Load the full pipeline with your custom tokenizer nlp = spacy.load("en_core_web_sm", tokenizer=tokenizer) # Test it doc = nlp("I don't like coffee. She isn't coming.") print([token.text for token in doc]) # Output: ["I", "do", "not", "like", "coffee", ".", "She", "is", "not", "coming", "."]
Why Your Original Approach Might Have Failed
Common pitfalls include:
- Forgetting to add your custom logic to the pipeline, or adding it too late (after the parser/tagger, so changes didn’t show up in your output).
- SpaCy’s built-in contraction rules overriding your custom rules—since the tokenizer applies default exceptions first, you need to either override them explicitly or use post-processing.
Either solution above should get you the result you want. Pick the one that fits your workflow best!
内容的提问来源于stack exchange,提问作者LCP

