You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为Spacy Tokenizer添加n't例外:n't转not失败的问题求解

Fixing SpaCy Tokenizer to Convert "n't" to "not"

Hey there! I get it—tweaking SpaCy's tokenizer to handle contractions exactly how you want can feel tricky. Let’s break down why your current approach might not be working and walk through two solid solutions to fix it.

First, a quick recap of SpaCy’s default behavior: The English tokenizer splits contractions like "don’t" into two separate tokens: do and n’t. If your attempts to replace "n’t" with "not" aren’t sticking, it’s likely because you’re either missing a step in the pipeline or not overriding SpaCy’s built-in contraction rules correctly.

Solution 1: Add a Custom Pipeline Component (Easiest & Most Flexible)

This approach modifies tokens right after they’re split, which is straightforward and doesn’t require messing with SpaCy’s core tokenizer logic. Here’s how to implement it:

import spacy

nlp = spacy.load("en_core_web_sm")

# Define a component to swap "n't" tokens with "not"
def replace_nt_with_not(doc):
    for token in doc:
        if token.text == "n't":
            token.text = "not"
            token.lemma_ = "not"  # Optional: Update lemma for consistency
    return doc

# Insert the component early in the pipeline (before parser/tagger)
nlp.add_pipe(replace_nt_with_not, before="parser")

# Test it out
doc = nlp("I don't like coffee. She isn't coming.")
print([token.text for token in doc])
# Output: ["I", "do", "not", "like", "coffee", ".", "She", "is", "not", "coming", "."]

This works because the component runs immediately after tokenization, so every "n’t" token gets replaced before any downstream processing (like tagging or parsing) happens.

Solution 2: Modify Tokenizer Exceptions (Change Split Behavior)

If you want SpaCy to never split "n’t" as a separate token and instead tokenize contractions like "don’t" directly into "do" + "not", you can override the tokenizer’s exception rules. This requires explicitly defining each contraction you want to handle:

import spacy
from spacy.lang.en import English
from spacy.symbols import ORTH

# Initialize a base tokenizer
nlp = English()
tokenizer = nlp.tokenizer

# Map contractions to their expanded token splits
contraction_mappings = {
    "don't": ["do", "not"],
    "isn't": ["is", "not"],
    "wasn't": ["was", "not"],
    "haven't": ["have", "not"],
    # Add all other contractions you need here
}

# Update the tokenizer's special cases
for contraction, split_tokens in contraction_mappings.items():
    tokenizer.add_special_case(contraction, [{ORTH: token} for token in split_tokens])

# Load the full pipeline with your custom tokenizer
nlp = spacy.load("en_core_web_sm", tokenizer=tokenizer)

# Test it
doc = nlp("I don't like coffee. She isn't coming.")
print([token.text for token in doc])
# Output: ["I", "do", "not", "like", "coffee", ".", "She", "is", "not", "coming", "."]

Why Your Original Approach Might Have Failed

Common pitfalls include:

  • Forgetting to add your custom logic to the pipeline, or adding it too late (after the parser/tagger, so changes didn’t show up in your output).
  • SpaCy’s built-in contraction rules overriding your custom rules—since the tokenizer applies default exceptions first, you need to either override them explicitly or use post-processing.

Either solution above should get you the result you want. Pick the one that fits your workflow best!

内容的提问来源于stack exchange,提问作者LCP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:11:40