正则表达式分词含数字问题:如何将带数字短语作为整体分词?
Hey there! Let's sort out this tokenization problem you're facing.
First, let's break down why your current code isn't working:
- Your regex pattern
r"[\w\s\d]+"matches all letters, numbers, and spaces as a single continuous sequence. That's why you're getting the entire sentence as one token instead of the split you want. Also,\walready includes digits (\dis a subset of\w), so[\w\s\d]is redundant—it's the same as[\w\s].
Based on your desired output ['this', 'is', 'an', 'example 123'], it looks like you want to treat "example 123" (with a space between the word and number) as a single token, while keeping other standalone words separate. Here's how to do that:
Solution for Combining Words + Space + Numbers as One Token
We'll use a regex pattern that matches either a standalone letter-only word, or a letter word followed by a space and digits:
from nltk.tokenize import RegexpTokenizer # Regex pattern: matches letter-only words OR words + space + digits token_pattern = r"\b[a-zA-Z]+(?:\s\d+)?\b" tokenizer = RegexpTokenizer(token_pattern) # Test it out result = tokenizer.tokenize("this is an example 123") print(result) # Output: ['this', 'is', 'an', 'example 123']
Let's break down the pattern:
\b: Word boundary to ensure we match full words, not partial ones[a-zA-Z]+: Matches one or more letters (the core word)(?:\s\d+)?: Optional non-capturing group that matches a space followed by one or more digits (the?makes this part optional)\b: Closing word boundary
Alternative: If You Mean Combined Word+Number (No Space)
If your actual goal is to treat combined terms like "example123" (no space between word and number) as a single token (instead of splitting into "example" and "123"), you can use a simpler pattern since \w already includes letters and digits:
from nltk.tokenize import RegexpTokenizer token_pattern = r"\w+" tokenizer = RegexpTokenizer(token_pattern) result = tokenizer.tokenize("this is an example123") print(result) # Output: ['this', 'is', 'an', 'example123']
This pattern matches any sequence of letters, digits, or underscores—perfect for keeping alphanumeric terms intact.
内容的提问来源于stack exchange,提问作者Maryam

