如何在Pandas DataFrame中提取指定值并移除关联列表对应元素
处理Pandas DataFrame中列表元素的移除与迁移
问题场景
现有一个Pandas DataFrame包含三列:
- A列:存储单个值
- B列:存储包含A列对应值的元组/列表
- C列:存储与B列元素一一对应的元组/列表
需要实现:从B列中移除A列的对应值,同时移除C列中对应位置的元素,并将这两个被移除的值分别存入新列D和E。
示例数据
当前DataFrame:
| A | B | C |
|---|---|---|
| 12 | (12, 20, 25) | (100, 200, 300) |
| 34 | (34, 3, 4, 5) | (10, 15, 20, 900) |
| 56 | (101, 56, 7) | (98, 97, 96) |
| 78 | (78) | (13) |
期望转换后的DataFrame:
| A | B | C | D | E |
|---|---|---|---|---|
| 12 | 20, 25 | 200, 300 | 12 | 100 |
| 34 | 3, 4, 5 | 15, 20, 900 | 34 | 10 |
| 56 | 101, 7 | 98, 96 | 56 | 97 |
| 78 | 78 | 13 |
解决方案
可以通过apply函数逐行处理数据,针对每行完成元素查找、移除与赋值操作:
import pandas as pd def process_row(row): val_a = row['A'] list_b = list(row['B']) list_c = list(row['C']) # 定位A值在B列表中的索引 idx = list_b.index(val_a) # 提取被移除的元素到D、E列 row['D'] = val_a row['E'] = list_c.pop(idx) # 移除B、C列中对应位置的元素 list_b.pop(idx) # 将剩余元素转为逗号分隔字符串(若需保留列表类型,直接赋值list_b/list_c即可) row['B'] = ', '.join(map(str, list_b)) if list_b else '' row['C'] = ', '.join(map(str, list_c)) if list_c else '' return row # 构建示例DataFrame df = pd.DataFrame({ 'A': [12, 34, 56, 78], 'B': [(12,20,25), (34,3,4,5), (101,56,7), (78,)], 'C': [(100,200,300), (10,15,20,900), (98,97,96), (13,)] }) # 应用处理函数 df = df.apply(process_row, axis=1) print(df)
异常处理(可选)
如果存在B列不包含A列值的情况,可以添加异常捕获逻辑,避免报错:
def process_row(row): val_a = row['A'] list_b = list(row['B']) list_c = list(row['C']) try: idx = list_b.index(val_a) except ValueError: # 找不到匹配值时,D、E列设为缺失值,B、C列保持原样 row['D'] = pd.NA row['E'] = pd.NA return row row['D'] = val_a row['E'] = list_c.pop(idx) list_b.pop(idx) row['B'] = ', '.join(map(str, list_b)) if list_b else '' row['C'] = ', '.join(map(str, list_c)) if list_c else '' return row
内容的提问来源于stack exchange,提问作者Thenelly
相关产品推荐
相关产品推荐

