如何高效删除DataFrame每行列表中指定索引的元素?
Got it, let's work through this problem together. You have a DataFrame where each row in the Text column is a list, and you need to keep only the elements at indices 2 and 4 (0-based) from each list—plus you want an efficient solution instead of slow loops.
First, let's break down why your previous attempts failed:
- Iterrows + del error: When you ran
for index,rows in df.iterrows(): del df.Text[0:3], you were trying to delete the first 3 rows of the entireTextcolumn, not elements inside each list. That's why you got the "invalid key" error for the slice. - Pop method error:
x.pop[0:3]is wrong becausepop()is a function, not a subscriptable object. You need to call it with parentheses, and it only accepts a single index (not a slice) anyway—so this approach wouldn't work for deleting multiple elements cleanly.
Efficient Solutions
Option 1: Directly Extract Desired Indices (Cleanest & Fastest)
The simplest and most efficient way is to use apply() with a lambda that grabs exactly the elements at indices 2 and 4 from each list:
import pandas as pd # Your original DataFrame df = pd.DataFrame({"Text": [[1,2,3,4,5], ['abc','bcd','cd','d','e'], ['a',1,'de',3,'fg']]}) # Extract elements at indices 2 and 4 df['Text'] = df['Text'].apply(lambda x: [x[2], x[4]])
This will give you exactly the output you want:
0 [3, 5] 1 ['cd', 'e'] 2 ['de', 'fg'] Name: Text, dtype: object
Option 2: Remove Unwanted Indices (If You Prefer Deletion)
If you specifically want to delete indices 0, 1, 3 instead of extracting, use a list comprehension inside apply() to filter out unwanted indices. This avoids issues with shifting indices when deleting elements in sequence:
df['Text'] = df['Text'].apply(lambda x: [val for idx, val in enumerate(x) if idx in (2, 4)])
Alternatively, if you really want to use deletion (note: this is less intuitive), you have to delete higher indices first to avoid shifting the positions of elements you still need to remove:
df['Text'] = df['Text'].apply(lambda x: (del x[3], del x[0:2], x)[2])
Why These Work
Both options use apply() which is vectorized and much faster than looping with iterrows() for large DataFrames. The first option is preferred because it's straightforward and avoids any risk of index-shifting mistakes from deletion.
内容的提问来源于stack exchange,提问作者user10395806

