如何快速从Pandas DataFrame中移除指定负面词汇
Iterating over DataFrame rows with itertuples() is definitely slow for large datasets—Pandas is designed to shine with vectorized operations instead. Here's a much faster approach that leverages Pandas' built-in string methods and set lookups (which are O(1) vs. O(n) for lists):
Step-by-Step Explanation & Code
1. Preprocess Negative Words for Fast Lookup
First, convert your negative words list into an uppercase set (since your GOODS_DESC values are uppercase). Set lookups are way faster than list checks, which is critical for performance with 1000+ negative words:
# Convert negative words to uppercase set for O(1) lookups negative_words_set = {word.upper() for word in negative_words_list}
2. Vectorized Processing of GOODS_DESC
Instead of looping through each row, use Pandas' string methods to split, filter, and rejoin the words in one go:
import pandas as pd # Sample DataFrame matching your structure data = { 'FileName': [17668620]*9, 'PageNo': ['TM000004', 'TM000014', 'TM000014', 'TM000014', 'TM000016', 'TM000016', 'TM000016', 'TM000016', 'TM000017'], 'LineNo': [36, 41, 42, 49, 49, 29, 32, 37, 113], 'GOODS_DESC': [ 'CAST ARTICLES IRON SANITARY', 'CRATES', 'CAST ARTICLES IRON', 'JAN ANIMAL AND VEGETABLE', 'SETTLING AGENT', 'JAN', 'CLAUSES SPECIAL CONDITIONS WARRANTIES', 'CARGO ISM ENDORSEMENT', 'QUANTITY DECLARED IRON CRATES' ] } df = pd.DataFrame(data) # Process GOODS_DESC: split into words, filter out negatives, rejoin df['GOODS_DESC'] = df['GOODS_DESC'].str.split() \ .apply(lambda x: [word for word in x if word not in negative_words_set]) \ .str.join(' ') # Replace empty strings with NaN to match your target output df['GOODS_DESC'] = df['GOODS_DESC'].replace('', pd.NA) # Optional: Drop rows where GOODS_DESC is NaN (uncomment if needed) # df = df.dropna(subset=['GOODS_DESC'])
3. Verify the Output
This will produce exactly the target DataFrame you shared:
| FileName | PageNo | LineNo | GOODS_DESC | |
|---|---|---|---|---|
| 0 | 17668620 | TM000004 | 36 | IRON |
| 1 | 17668620 | TM000014 | 41 | CRATES |
| 2 | 17668620 | TM000014 | 42 | IRON |
| 3 | 17668620 | TM000014 | 49 | ANIMAL AND VEGETABLE |
| 4 | 17668620 | TM000016 | 49 | SETTLING AGENT |
| 5 | 17668620 | TM000016 | 29 | |
| 6 | 17668620 | TM000016 | 32 | CONDITIONS WARRANTIES |
| 7 | 17668620 | TM000016 | 37 | CARGO ISM |
| 8 | 17668620 | TM000017 | 113 | QUANTITY DECLARED IRON CRATES |
Why This Is Faster
- Set Lookups: Checking if a word is in a set takes constant time (
O(1)), compared toO(n)for a list (wherenis 1000 in your case). - Vectorized Operations: Pandas string methods (
str.split,str.join) andapplyon Series are optimized under the hood (using C extensions), so they avoid the overhead of Python-level loops over rows. - No In-Place Modifications: The original code uses
df.locanddropin-place, which can be slow. This approach processes the entire column in one pass.
For even larger datasets, you could use str.replace with a regex pattern built from the negative words, but the above method is more readable and still extremely fast for most cases.
内容的提问来源于stack exchange,提问作者Madhur Yadav

