You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效移除文本中的非字母数字字符与停用词?解决Pandas DataFrame赋值时的长度不匹配问题

Answers to Your Pandas + NLTK Questions

Issue 2: ValueError When Adding Tokens to DataFrame

First, let's fix the length mismatch error. The error ValueError: Length of values (508) does not match length of index (251) almost always means your tokens_linhas list has more elements than your DataFrame has rows.

The most common cause here is not resetting tokens_linhas to an empty list before running the loop with stopword removal. If you first ran the loop without stopword checks (which added 251 elements) and then ran the stopword-filtered loop without clearing tokens_linhas, you’d end up with 251 + 251 = 508 elements.

Quick fix: Always initialize tokens_linhas right before your loop:

tokens_linhas = []  # Reset to empty list first!
for index, value in df.text.iteritems():
    tokens = tokenize.word_tokenize(value)
    tokens_sem_pont = [token for token in tokens if token.isalnum() and (token not in stopwords.words('english'))]
    tokens_linhas.append(tokens_sem_pont)

If you still see the error after this, double-check that your DataFrame doesn’t have duplicate rows or that df.text.iteritems() is indeed iterating exactly once per row (you can verify with print(len(df)) vs print(len(tokens_linhas)) mid-loop).

Issue 1: Optimizing Slow Execution

Your original code is slow mainly because of two bottlenecks:

  1. Checking membership in a list: stopwords.words('english') returns a list, and token not in list is an O(n) operation (slow for large numbers of tokens).
  2. Manual for loop: Iterating with iteritems() is less efficient than using pandas’ vectorized operations.

Here’s how to fix both:

Step 1: Precompute Stopwords as a Set

Sets have O(1) membership checks, which is way faster. Do this once outside your loop/function:

from nltk.corpus import stopwords
stop_words = set(stopwords.words('english'))  # Convert to set once

Step 2: Use Pandas apply() Instead of a Manual Loop

Pandas’ apply() is optimized for column-wise operations and will run faster than a manual for loop. Combine tokenization and filtering into a reusable function:

import pandas as pd
from nltk.tokenize import word_tokenize

def process_text(text):
    # Handle NaN values to avoid errors
    if pd.isna(text):
        return []
    # Tokenize and filter in one step
    tokens = word_tokenize(text)
    return [token for token in tokens if token.isalnum() and token not in stop_words]

# Apply the function to your text column directly
df['tokens'] = df['text'].apply(process_text)

Bonus: Even Faster Options (For Large Datasets)

If you’re working with a huge dataset, consider these additional optimizations:

  • Parallel Processing: Use the swifter library to automatically parallelize the apply() call:
    import swifter
    df['tokens'] = df['text'].swifter.apply(process_text)
    
  • Faster Tokenizers: Replace NLTK’s word_tokenize with spaCy’s tokenizer (which is faster for large volumes) or sklearn’s CountVectorizer (though it’s designed for bag-of-words, you can adapt it to get tokens).

内容的提问来源于stack exchange,提问作者Pedro Kaneto Suzuki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 16:52:49