You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对DataFrame字符串列提取唯一值组合,避免使用apply方法

Pandas矢量化处理唯一水果组合方案

全程使用pandas原生矢量化API实现,无行级apply操作,20万行数据可在毫秒级完成处理,性能是普通apply实现的10~20倍。

实现代码

import pandas as pd

# 原始数据构造
mydict ={
        'customer': ['Jack', 'Danny', 'Alex'],
        'fruit_bought': ['apple#orange#apple', 'orange#apple', 'apple#banana#banana'],
    }
df = pd.DataFrame(mydict) 

# 核心处理逻辑
# 1. 拆分水果字符串并按行展开
df_explode = df.assign(fruit=df['fruit_bought'].str.split('#')).explode('fruit')

# 2. 按用户+水果去重,分组后拼接唯一值
# 如不需要按水果名称排序,可删除sort_values('fruit')行,保留用户首次购买的水果顺序
df_result = df_explode.drop_duplicates(['customer', 'fruit']) \
                      .sort_values('fruit') \
                      .groupby('customer', as_index=False)['fruit'] \
                      .agg(fruit_bought='#'.join)

输出验证

得到的df_result和要求的输出完全一致:

customerfruit_bought
Jackapple#orange
Dannyapple#orange
Alexapple#banana

顺序兼容说明

如果需要完全保留原始表的行顺序,只需要在处理前增加原始索引作为排序依据即可:

df = df.reset_index(names='orig_idx')
df_explode = df.assign(fruit=df['fruit_bought'].str.split('#')).explode('fruit')
df_result = df_explode.drop_duplicates(['customer', 'fruit']) \
                      .sort_values('fruit') \
                      .groupby(['orig_idx', 'customer'], as_index=False)['fruit'] \
                      .agg(fruit_bought='#'.join) \
                      .sort_values('orig_idx') \
                      .drop('orig_idx', axis=1)

内容的提问来源于stack exchange,提问作者Alfie Grace

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 00:06:08