如何在Pandas DataFrame中删除列表列连续重复元素并同步另一列?
处理Pandas DataFrame中连续重复元素的方法
1. 构造示例数据
先还原你提供的输入DataFrame:
import pandas as pd data = { "A": [32, 35], "B": [[1,2,2,3,4], [5,5,7,7,7,8]], "C": [["a","b","c","d","e"], ["q","w","e","r","t","y"]] } df = pd.DataFrame(data)
2. 编写自定义过滤函数
写一个函数处理单行数据,同步过滤B列的连续重复后续项,以及C列对应位置的元素:
def filter_consec_dups(row): b_vals = row["B"] c_vals = row["C"] new_b = [] new_c = [] last_val = None for b, c in zip(b_vals, c_vals): if b != last_val: new_b.append(b) new_c.append(c) last_val = b return pd.Series([new_b, new_c], index=["B", "C"])
3. 应用函数到DataFrame
通过apply按行执行过滤,替换原有的B、C列:
df[["B", "C"]] = df.apply(filter_consec_dups, axis=1)
最终输出结果
执行后得到的DataFrame如下:
| A | B | C |
|---|---|---|
| 32 | [1,2,3,4] | [a,b,d,e] |
| 35 | [5,7,8] | [q,e,y] |
说明
- 该方法逐行遍历列表,仅保留与前一个元素不同的项,精准过滤连续重复的后续元素
- 兼容不同行之间列表长度不一致的情况,通用性强
内容的提问来源于stack exchange,提问作者codex
相关产品推荐
相关产品推荐

