如何对DataFrame字符串列提取唯一值组合,避免使用apply方法
Pandas矢量化处理唯一水果组合方案
全程使用pandas原生矢量化API实现,无行级apply操作,20万行数据可在毫秒级完成处理,性能是普通apply实现的10~20倍。
实现代码
import pandas as pd # 原始数据构造 mydict ={ 'customer': ['Jack', 'Danny', 'Alex'], 'fruit_bought': ['apple#orange#apple', 'orange#apple', 'apple#banana#banana'], } df = pd.DataFrame(mydict) # 核心处理逻辑 # 1. 拆分水果字符串并按行展开 df_explode = df.assign(fruit=df['fruit_bought'].str.split('#')).explode('fruit') # 2. 按用户+水果去重,分组后拼接唯一值 # 如不需要按水果名称排序,可删除sort_values('fruit')行,保留用户首次购买的水果顺序 df_result = df_explode.drop_duplicates(['customer', 'fruit']) \ .sort_values('fruit') \ .groupby('customer', as_index=False)['fruit'] \ .agg(fruit_bought='#'.join)
输出验证
得到的df_result和要求的输出完全一致:
| customer | fruit_bought |
|---|---|
| Jack | apple#orange |
| Danny | apple#orange |
| Alex | apple#banana |
顺序兼容说明
如果需要完全保留原始表的行顺序,只需要在处理前增加原始索引作为排序依据即可:
df = df.reset_index(names='orig_idx') df_explode = df.assign(fruit=df['fruit_bought'].str.split('#')).explode('fruit') df_result = df_explode.drop_duplicates(['customer', 'fruit']) \ .sort_values('fruit') \ .groupby(['orig_idx', 'customer'], as_index=False)['fruit'] \ .agg(fruit_bought='#'.join) \ .sort_values('orig_idx') \ .drop('orig_idx', axis=1)
内容的提问来源于stack exchange,提问作者Alfie Grace
相关产品推荐
相关产品推荐

