You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式分词含数字问题:如何将带数字短语作为整体分词?

Fixing Tokenization to Include Numbers with Words

Hey there! Let's sort out this tokenization problem you're facing.

First, let's break down why your current code isn't working:

  • Your regex pattern r"[\w\s\d]+" matches all letters, numbers, and spaces as a single continuous sequence. That's why you're getting the entire sentence as one token instead of the split you want. Also, \w already includes digits (\d is a subset of \w), so [\w\s\d] is redundant—it's the same as [\w\s].

Based on your desired output ['this', 'is', 'an', 'example 123'], it looks like you want to treat "example 123" (with a space between the word and number) as a single token, while keeping other standalone words separate. Here's how to do that:

Solution for Combining Words + Space + Numbers as One Token

We'll use a regex pattern that matches either a standalone letter-only word, or a letter word followed by a space and digits:

from nltk.tokenize import RegexpTokenizer

# Regex pattern: matches letter-only words OR words + space + digits
token_pattern = r"\b[a-zA-Z]+(?:\s\d+)?\b"
tokenizer = RegexpTokenizer(token_pattern)

# Test it out
result = tokenizer.tokenize("this is an example 123")
print(result)  # Output: ['this', 'is', 'an', 'example 123']

Let's break down the pattern:

  • \b: Word boundary to ensure we match full words, not partial ones
  • [a-zA-Z]+: Matches one or more letters (the core word)
  • (?:\s\d+)?: Optional non-capturing group that matches a space followed by one or more digits (the ? makes this part optional)
  • \b: Closing word boundary

Alternative: If You Mean Combined Word+Number (No Space)

If your actual goal is to treat combined terms like "example123" (no space between word and number) as a single token (instead of splitting into "example" and "123"), you can use a simpler pattern since \w already includes letters and digits:

from nltk.tokenize import RegexpTokenizer

token_pattern = r"\w+"
tokenizer = RegexpTokenizer(token_pattern)

result = tokenizer.tokenize("this is an example123")
print(result)  # Output: ['this', 'is', 'an', 'example123']

This pattern matches any sequence of letters, digits, or underscores—perfect for keeping alphanumeric terms intact.

内容的提问来源于stack exchange,提问作者Maryam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:00:07