You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

分组后如何去除Pandas DataFrame中的重复字符串?

Pandas按ID分组并去重字符串列的解决方案

原始数据

IDElementsColors
A'1st element, 2d element, 3d element''red, blue'
A'2d element, 4th element''blue, green'
B'3d element, 5th element, 6th element''white, purple'
B'3d element, 5th element, 7th element''white, teal'
B'3d element, 5th element, 8th element''white, black'

期望结果

IDElementsColors
A'1st element, 2d element, 3d element, 4th element''red, blue, green'
B'3d element, 5th element, 6th element, 7th element, 8th element''white, purple, teal, black'

实现代码

1. 构造初始DataFrame

import pandas as pd

data = {
    'ID': ['A', 'A', 'B', 'B', 'B'],
    'Elements': [
        "'1st element, 2d element, 3d element'",
        "'2d element, 4th element'",
        "'3d element, 5th element, 6th element'",
        "'3d element, 5th element, 7th element'",
        "'3d element, 5th element, 8th element'"
    ],
    'Colors': [
        "'red, blue'",
        "'blue, green'",
        "'white, purple'",
        "'white, teal'",
        "'white, black'"
    ]
}

df = pd.DataFrame(data)

2. 定义字符串去重函数

这个函数负责处理单引号包裹的字符串,拆分元素、去重后重新拼接:

def deduplicate_string(s):
    # 移除首尾单引号
    s_clean = s.strip("'")
    # 拆分元素并去除每个元素的前后空格
    items = [item.strip() for item in s_clean.split(',')]
    # 去重并保留原始出现顺序
    unique_items = list(dict.fromkeys(items))
    # 重新拼接成带单引号的字符串
    return f"'{', '.join(unique_items)}'"

3. 分组并应用去重逻辑

按ID分组,对Elements和Colors列分别应用去重函数:

result_df = df.groupby('ID').agg({
    'Elements': lambda x: deduplicate_string(', '.join(x)),
    'Colors': lambda x: deduplicate_string(', '.join(x))
}).reset_index()

4. 查看最终结果

print(result_df)

运行后即可得到符合期望的DataFrame,其中dict.fromkeys的使用保证了元素的原始出现顺序,避免了用set去重导致的顺序混乱问题。

内容的提问来源于stack exchange,提问作者crocefisso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 04:35:45