如何高效移除文本中的非字母数字字符与停用词?解决Pandas DataFrame赋值时的长度不匹配问题
Issue 2: ValueError When Adding Tokens to DataFrame
First, let's fix the length mismatch error. The error ValueError: Length of values (508) does not match length of index (251) almost always means your tokens_linhas list has more elements than your DataFrame has rows.
The most common cause here is not resetting tokens_linhas to an empty list before running the loop with stopword removal. If you first ran the loop without stopword checks (which added 251 elements) and then ran the stopword-filtered loop without clearing tokens_linhas, you’d end up with 251 + 251 = 508 elements.
Quick fix: Always initialize tokens_linhas right before your loop:
tokens_linhas = [] # Reset to empty list first! for index, value in df.text.iteritems(): tokens = tokenize.word_tokenize(value) tokens_sem_pont = [token for token in tokens if token.isalnum() and (token not in stopwords.words('english'))] tokens_linhas.append(tokens_sem_pont)
If you still see the error after this, double-check that your DataFrame doesn’t have duplicate rows or that df.text.iteritems() is indeed iterating exactly once per row (you can verify with print(len(df)) vs print(len(tokens_linhas)) mid-loop).
Issue 1: Optimizing Slow Execution
Your original code is slow mainly because of two bottlenecks:
- Checking membership in a list:
stopwords.words('english')returns a list, andtoken not in listis an O(n) operation (slow for large numbers of tokens). - Manual for loop: Iterating with
iteritems()is less efficient than using pandas’ vectorized operations.
Here’s how to fix both:
Step 1: Precompute Stopwords as a Set
Sets have O(1) membership checks, which is way faster. Do this once outside your loop/function:
from nltk.corpus import stopwords stop_words = set(stopwords.words('english')) # Convert to set once
Step 2: Use Pandas apply() Instead of a Manual Loop
Pandas’ apply() is optimized for column-wise operations and will run faster than a manual for loop. Combine tokenization and filtering into a reusable function:
import pandas as pd from nltk.tokenize import word_tokenize def process_text(text): # Handle NaN values to avoid errors if pd.isna(text): return [] # Tokenize and filter in one step tokens = word_tokenize(text) return [token for token in tokens if token.isalnum() and token not in stop_words] # Apply the function to your text column directly df['tokens'] = df['text'].apply(process_text)
Bonus: Even Faster Options (For Large Datasets)
If you’re working with a huge dataset, consider these additional optimizations:
- Parallel Processing: Use the
swifterlibrary to automatically parallelize theapply()call:import swifter df['tokens'] = df['text'].swifter.apply(process_text) - Faster Tokenizers: Replace NLTK’s
word_tokenizewith spaCy’s tokenizer (which is faster for large volumes) or sklearn’sCountVectorizer(though it’s designed for bag-of-words, you can adapt it to get tokens).
内容的提问来源于stack exchange,提问作者Pedro Kaneto Suzuki

