SpaCy如何将词内连字符视为单个词?相关正则作用解析
Understanding and Fixing SpaCy Tokenization for Hyphenated Words
Let's break down your questions about the regex patterns in the custom tokenizer, then walk through how to achieve your desired tokenization result.
Explanation of the Regex Patterns in Your Original Infixes
1. r"[./]"
This is a character class that matches either a dot (.) or forward slash (/). In SpaCy's tokenization system, infix patterns define sequences inside a token that should split it into multiple tokens.
- In practice, this pattern would split tokens at standalone dots or slashes (e.g., turning "file.txt" into ["file", ".", "txt"]), but it doesn't affect numeric values like "3.14"—the default SpaCy rules already preserve dots between digits as part of a single token.
2. r"(.'.')"
This regex is likely a typo, but as written, it matches a very specific sequence: .'.' (dot, apostrophe, dot, apostrophe).
- The bigger issue here is that adding this (along with mixing prefixes into infixes) broke the default handling of apostrophes. In your test case, "Yahya's" split into ["Yahya", "'", "s"] instead of the desired ["Yahya", "'s"] because the custom infix rules interfered with SpaCy's default prefix logic for contractions.
Fixing the Tokenization to Match Your Desired Output
Your goal is to:
- Keep hyphenated words (like
laptop-cover) as single tokens - Preserve default contraction handling (e.g.,
Yahya's→ ["Yahya", "'s"]) - Keep numeric values (like
3.14) intact - Split trailing punctuation (like
.afterlaptop-cover) as separate tokens
Here's the corrected approach:
Corrected Code Example
import spacy from spacy.tokenizer import Tokenizer from spacy.util import compile_infix_regex nlp = spacy.load('en_core_web_sm') # Get default infixes and remove the hyphen-splitting pattern default_infixes = list(nlp.Defaults.infixes) # Filter out the regex that splits tokens at hyphens default_infixes = [infix for infix in default_infixes if not (infix.startswith(r'(?<=[\w])[') and '-' in infix)] # Compile the modified infix regex infix_re = compile_infix_regex(default_infixes) def custom_tokenizer(nlp): return Tokenizer( nlp.vocab, prefix_search=nlp.tokenizer.prefix_search, # Keep default prefix rules (for apostrophes) suffix_search=nlp.tokenizer.suffix_search, # Keep default suffix rules (for punctuation) infix_finditer=infix_re.finditer, token_match=nlp.tokenizer.token_match ) nlp.tokenizer = custom_tokenizer(nlp) # Test your target sentence test_sentence = "That is Yahya's laptop-cover. 3.14!" doc = nlp(test_sentence) print([token.text for token in doc]) # Output: ['That', 'is', 'Yahya', "'s", 'laptop-cover', '.', '3.14', '!'] # Test additional examples s1 = "Marketing-Representative- won't die in car accident." s2 = "Out-of-box implementation" for s in [s1, s2]: doc = nlp(s) print([token.text for token in doc]) # Output for s1: ['Marketing-Representative-', 'wo', "n't", 'die', 'in', 'car', 'accident', '.'] # Output for s2: ['Out-of-box', 'implementation']
Why This Works
- Hyphenated words: By removing the hyphen pattern from infixes, SpaCy no longer splits tokens at hyphens, keeping words like
laptop-coverintact. - Apostrophe handling: We retained SpaCy's default prefix rules, so contractions like
'sare correctly split as separate tokens. - Numeric values: The default rules preserving dots between digits remain active, so
3.14stays as a single token. - Punctuation: Default suffix rules split trailing punctuation (like
.and!) into separate tokens, matching your desired output.
What Was Wrong with the Original Code
- Mixing prefix patterns into the infixes list broke the apostrophe contraction logic.
- The
r"[-]~"pattern merged hyphens with surrounding characters but also caused trailing punctuation to stick to hyphenated words (e.g.,laptop-cover.instead of separatelaptop-coverand.).
内容的提问来源于stack exchange,提问作者Zeeshan Ali
相关产品推荐
相关产品推荐

