如何将DataFrame指定列的竖线分隔值拆分为多行并对结果去重?
高效实现方案
直接使用Pandas内置的矢量化方法即可实现需求,相比手动遍历iterrows的方式性能提升显著,代码也更简洁:
核心逻辑
- 用
str.split('|')将目标列的字符串按分隔符拆分为列表 - 用
explode()方法将列表拆分为多行,其余列值自动保持与原行一致 - 用
drop_duplicates()直接完成全量去重
完整代码示例
import pandas as pd processed_dfs = [] for df in all_dfs: # 不包含待拆分id列的DataFrame直接保留 if 'id' not in df.columns: processed_dfs.append(df) continue # 链式调用完成拆分、展开、去重全流程 processed_df = df.assign(id=df['id'].str.split('|')) \ .explode('id', ignore_index=True) \ .drop_duplicates(ignore_index=True) processed_dfs.append(processed_df) # 若需要将所有处理后的DataFrame合并为一个总表,可增加以下代码 # total_df = pd.concat(processed_dfs, ignore_index=True).drop_duplicates(ignore_index=True)
补充说明
- 如果需要过滤掉id列空值产生的无效行,可以在链式调用中增加
.dropna(subset=['id'])步骤 drop_duplicates默认根据所有列的值判断重复,若需要指定按特定列去重,可传入subset参数,例如drop_duplicates(subset=['id', 'uid'], ignore_index=True)- 该方案全程使用Pandas内置优化方法,数据量越大相对手动遍历的性能优势越明显
内容的提问来源于stack exchange,提问作者marlon
相关产品推荐
相关产品推荐

