You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于指定词表实现自定义分词:将特定短语设为单个token

How to Keep Specific Phrases as Single Tokens with NLTK's word_tokenize

Absolutely doable! Let me walk you through two reliable ways to make NLTK's word_tokenize treat your specific phrases ("nlp - nltk" and "CIFA R12 - INV") as single tokens instead of splitting them apart.

Method 1: Preprocess with Placeholders (Simple & Straightforward)

This approach works by temporarily replacing your target phrases with unique, non-splittable placeholders before tokenizing, then swapping them back afterward. It’s easy to maintain even if you add more phrases later.

Here’s the code example:

from nltk.tokenize import word_tokenize

# Map your target phrases to unique placeholders
phrase_placeholder_map = {
    "nlp - nltk": "NLPNLTK_UNIQUE_TOKEN",
    "CIFA R12 - INV": "CIFAR12INV_UNIQUE_TOKEN"
}
# Reverse the map for later replacement back to original phrases
placeholder_phrase_map = {v: k for k, v in phrase_placeholder_map.items()}

input_text = "This is sample text for nlp - nltk CIFA R12 - INV ."

# Step 1: Replace phrases with placeholders
processed_text = input_text
for phrase, placeholder in phrase_placeholder_map.items():
    processed_text = processed_text.replace(phrase, placeholder)

# Step 2: Run standard word_tokenize
tokens = word_tokenize(processed_text)

# Step 3: Swap placeholders back to original phrases
final_tokens = [
    placeholder_phrase_map[token] if token in placeholder_phrase_map else token
    for token in tokens
]

print(final_tokens)
# Output: ['This', 'is', 'sample', 'text', 'for', 'nlp - nltk', 'CIFA R12 - INV', '.']

Method 2: Use RegexpTokenizer (Direct Custom Tokenization)

If you prefer to avoid placeholder replacements, you can build a custom tokenizer with RegexpTokenizer that prioritizes matching your target phrases as whole tokens, then falls back to standard word patterns.

Here’s how to implement it:

from nltk.tokenize import RegexpTokenizer
import re

# List of your target phrases
target_phrases = ["nlp - nltk", "CIFA R12 - INV"]
# Escape special characters (like hyphens) to avoid regex errors
escaped_phrases = [re.escape(phrase) for phrase in target_phrases]

# Build regex pattern: first match any target phrase, then match words/punctuation
token_pattern = f'({"|".join(escaped_phrases)})|\\w+|[^\w\\s]'

# Initialize the custom tokenizer
custom_tokenizer = RegexpTokenizer(token_pattern)

input_text = "This is sample text for nlp - nltk CIFA R12 - INV ."
tokens = custom_tokenizer.tokenize(input_text)

# Clean up any empty tokens from edge cases
final_tokens = [token.strip() for token in tokens if token.strip()]

print(final_tokens)
# Output: ['This', 'is', 'sample', 'text', 'for', 'nlp - nltk', 'CIFA R12 - INV', '.']

Quick Notes

  • Method 1 is great for small to medium phrase lists and requires minimal regex knowledge.
  • Method 2 is more efficient for larger phrase sets and integrates the tokenization logic into one step, but you need to ensure special characters in phrases are properly escaped.

内容的提问来源于stack exchange,提问作者Vignesh Muthu.S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:45:27