基于指定词表实现自定义分词:将特定短语设为单个token
Absolutely doable! Let me walk you through two reliable ways to make NLTK's word_tokenize treat your specific phrases ("nlp - nltk" and "CIFA R12 - INV") as single tokens instead of splitting them apart.
Method 1: Preprocess with Placeholders (Simple & Straightforward)
This approach works by temporarily replacing your target phrases with unique, non-splittable placeholders before tokenizing, then swapping them back afterward. It’s easy to maintain even if you add more phrases later.
Here’s the code example:
from nltk.tokenize import word_tokenize # Map your target phrases to unique placeholders phrase_placeholder_map = { "nlp - nltk": "NLPNLTK_UNIQUE_TOKEN", "CIFA R12 - INV": "CIFAR12INV_UNIQUE_TOKEN" } # Reverse the map for later replacement back to original phrases placeholder_phrase_map = {v: k for k, v in phrase_placeholder_map.items()} input_text = "This is sample text for nlp - nltk CIFA R12 - INV ." # Step 1: Replace phrases with placeholders processed_text = input_text for phrase, placeholder in phrase_placeholder_map.items(): processed_text = processed_text.replace(phrase, placeholder) # Step 2: Run standard word_tokenize tokens = word_tokenize(processed_text) # Step 3: Swap placeholders back to original phrases final_tokens = [ placeholder_phrase_map[token] if token in placeholder_phrase_map else token for token in tokens ] print(final_tokens) # Output: ['This', 'is', 'sample', 'text', 'for', 'nlp - nltk', 'CIFA R12 - INV', '.']
Method 2: Use RegexpTokenizer (Direct Custom Tokenization)
If you prefer to avoid placeholder replacements, you can build a custom tokenizer with RegexpTokenizer that prioritizes matching your target phrases as whole tokens, then falls back to standard word patterns.
Here’s how to implement it:
from nltk.tokenize import RegexpTokenizer import re # List of your target phrases target_phrases = ["nlp - nltk", "CIFA R12 - INV"] # Escape special characters (like hyphens) to avoid regex errors escaped_phrases = [re.escape(phrase) for phrase in target_phrases] # Build regex pattern: first match any target phrase, then match words/punctuation token_pattern = f'({"|".join(escaped_phrases)})|\\w+|[^\w\\s]' # Initialize the custom tokenizer custom_tokenizer = RegexpTokenizer(token_pattern) input_text = "This is sample text for nlp - nltk CIFA R12 - INV ." tokens = custom_tokenizer.tokenize(input_text) # Clean up any empty tokens from edge cases final_tokens = [token.strip() for token in tokens if token.strip()] print(final_tokens) # Output: ['This', 'is', 'sample', 'text', 'for', 'nlp - nltk', 'CIFA R12 - INV', '.']
Quick Notes
- Method 1 is great for small to medium phrase lists and requires minimal regex knowledge.
- Method 2 is more efficient for larger phrase sets and integrates the tokenization logic into one step, but you need to ensure special characters in phrases are properly escaped.
内容的提问来源于stack exchange,提问作者Vignesh Muthu.S

