如何修改CountVectorizer正则表达式以去除数字和下划线
Got it, let's fix this regex issue for your CountVectorizer! The problem with your current token_pattern is that \w matches letters, numbers, and underscores—which is exactly why those unwanted tokens are showing up. Here's how to adjust it:
Step 1: Understand the Fix
We need to restrict the regex to only match Unicode letters (including Spanish-specific characters like ñ, á, é, etc.) and exclude digits/underscores entirely.
Step 2: Modified Token Pattern
Replace your existing token_pattern with one of these options, depending on how strict you want to be:
Option 1: Match Any Unicode Letter (Recommended for Spanish)
This covers all Spanish alphabet characters (including accents and ñ) without hardcoding every possible character:
r'(?u)\b[\p{L}]{2,}\b'
Breakdown:
(?u): Enables Unicode mode (critical for handling Spanish special characters)\b: Word boundary to ensure we're matching full words[\p{L}]: Matches any Unicode letter (coversa-z,A-Z,áéíóú,ñ, etc.){2,}: Requires at least 2 characters (matches your original logic of\w\w+)
Option 2: Explicitly Match Spanish Characters
If you want to strictly limit to only Spanish alphabet characters (and exclude other language letters like ä or ê), use this:
r'(?u)\b[a-zA-ZáéíóúÁÉÍÓÚñÑ]{2,}\b'
Step 3: Full Implementation Code
Here's how to plug this into your CountVectorizer:
from sklearn.feature_extraction.text import CountVectorizer import nltk from nltk.corpus import stopwords # Download Spanish stopwords if you haven't already nltk.download('stopwords') # Use the modified token pattern vectorizer = CountVectorizer( stop_words=stopwords.words('spanish'), token_pattern=r'(?u)\b[\p{L}]{2,}\b' ) # Fit to your text data and get clean features # vectorizer.fit(your_text_corpus) # clean_features = vectorizer.get_feature_names_out()
Result
After this change, tokens like 000, _aitor, 2t0s6dgxnm will be filtered out entirely—only valid Spanish words (with no digits/underscores) will remain in your feature list.
内容的提问来源于stack exchange,提问作者ambigus9

