You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何快速从Pandas DataFrame中移除指定负面词汇

Efficient Pandas Solution to Remove Negative Words from DataFrame

Iterating over DataFrame rows with itertuples() is definitely slow for large datasets—Pandas is designed to shine with vectorized operations instead. Here's a much faster approach that leverages Pandas' built-in string methods and set lookups (which are O(1) vs. O(n) for lists):

Step-by-Step Explanation & Code

1. Preprocess Negative Words for Fast Lookup

First, convert your negative words list into an uppercase set (since your GOODS_DESC values are uppercase). Set lookups are way faster than list checks, which is critical for performance with 1000+ negative words:

# Convert negative words to uppercase set for O(1) lookups
negative_words_set = {word.upper() for word in negative_words_list}

2. Vectorized Processing of GOODS_DESC

Instead of looping through each row, use Pandas' string methods to split, filter, and rejoin the words in one go:

import pandas as pd

# Sample DataFrame matching your structure
data = {
    'FileName': [17668620]*9,
    'PageNo': ['TM000004', 'TM000014', 'TM000014', 'TM000014', 'TM000016', 'TM000016', 'TM000016', 'TM000016', 'TM000017'],
    'LineNo': [36, 41, 42, 49, 49, 29, 32, 37, 113],
    'GOODS_DESC': [
        'CAST ARTICLES IRON SANITARY',
        'CRATES',
        'CAST ARTICLES IRON',
        'JAN ANIMAL AND VEGETABLE',
        'SETTLING AGENT',
        'JAN',
        'CLAUSES SPECIAL CONDITIONS WARRANTIES',
        'CARGO ISM ENDORSEMENT',
        'QUANTITY DECLARED IRON CRATES'
    ]
}
df = pd.DataFrame(data)

# Process GOODS_DESC: split into words, filter out negatives, rejoin
df['GOODS_DESC'] = df['GOODS_DESC'].str.split() \
    .apply(lambda x: [word for word in x if word not in negative_words_set]) \
    .str.join(' ')

# Replace empty strings with NaN to match your target output
df['GOODS_DESC'] = df['GOODS_DESC'].replace('', pd.NA)

# Optional: Drop rows where GOODS_DESC is NaN (uncomment if needed)
# df = df.dropna(subset=['GOODS_DESC'])

3. Verify the Output

This will produce exactly the target DataFrame you shared:

FileNamePageNoLineNoGOODS_DESC
017668620TM00000436IRON
117668620TM00001441CRATES
217668620TM00001442IRON
317668620TM00001449ANIMAL AND VEGETABLE
417668620TM00001649SETTLING AGENT
517668620TM00001629
617668620TM00001632CONDITIONS WARRANTIES
717668620TM00001637CARGO ISM
817668620TM000017113QUANTITY DECLARED IRON CRATES

Why This Is Faster

  • Set Lookups: Checking if a word is in a set takes constant time (O(1)), compared to O(n) for a list (where n is 1000 in your case).
  • Vectorized Operations: Pandas string methods (str.split, str.join) and apply on Series are optimized under the hood (using C extensions), so they avoid the overhead of Python-level loops over rows.
  • No In-Place Modifications: The original code uses df.loc and drop in-place, which can be slow. This approach processes the entire column in one pass.

For even larger datasets, you could use str.replace with a regex pattern built from the negative words, but the above method is more readable and still extremely fast for most cases.

内容的提问来源于stack exchange,提问作者Madhur Yadav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:59:35