You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改CountVectorizer正则表达式以去除数字和下划线

Got it, let's fix this regex issue for your CountVectorizer! The problem with your current token_pattern is that \w matches letters, numbers, and underscores—which is exactly why those unwanted tokens are showing up. Here's how to adjust it:

Step 1: Understand the Fix

We need to restrict the regex to only match Unicode letters (including Spanish-specific characters like ñ, á, é, etc.) and exclude digits/underscores entirely.

Step 2: Modified Token Pattern

Replace your existing token_pattern with one of these options, depending on how strict you want to be:

Option 1: Match Any Unicode Letter (Recommended for Spanish)

This covers all Spanish alphabet characters (including accents and ñ) without hardcoding every possible character:

r'(?u)\b[\p{L}]{2,}\b'

Breakdown:

  • (?u): Enables Unicode mode (critical for handling Spanish special characters)
  • \b: Word boundary to ensure we're matching full words
  • [\p{L}]: Matches any Unicode letter (covers a-z, A-Z, áéíóú, ñ, etc.)
  • {2,}: Requires at least 2 characters (matches your original logic of \w\w+)

Option 2: Explicitly Match Spanish Characters

If you want to strictly limit to only Spanish alphabet characters (and exclude other language letters like ä or ê), use this:

r'(?u)\b[a-zA-ZáéíóúÁÉÍÓÚñÑ]{2,}\b'

Step 3: Full Implementation Code

Here's how to plug this into your CountVectorizer:

from sklearn.feature_extraction.text import CountVectorizer
import nltk
from nltk.corpus import stopwords

# Download Spanish stopwords if you haven't already
nltk.download('stopwords')

# Use the modified token pattern
vectorizer = CountVectorizer(
    stop_words=stopwords.words('spanish'),
    token_pattern=r'(?u)\b[\p{L}]{2,}\b'
)

# Fit to your text data and get clean features
# vectorizer.fit(your_text_corpus)
# clean_features = vectorizer.get_feature_names_out()

Result

After this change, tokens like 000, _aitor, 2t0s6dgxnm will be filtered out entirely—only valid Spanish words (with no digits/underscores) will remain in your feature list.

内容的提问来源于stack exchange,提问作者ambigus9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:57:07