分组后如何去除Pandas DataFrame中的重复字符串?
Pandas按ID分组并去重字符串列的解决方案
原始数据
| ID | Elements | Colors |
|---|---|---|
| A | '1st element, 2d element, 3d element' | 'red, blue' |
| A | '2d element, 4th element' | 'blue, green' |
| B | '3d element, 5th element, 6th element' | 'white, purple' |
| B | '3d element, 5th element, 7th element' | 'white, teal' |
| B | '3d element, 5th element, 8th element' | 'white, black' |
期望结果
| ID | Elements | Colors |
|---|---|---|
| A | '1st element, 2d element, 3d element, 4th element' | 'red, blue, green' |
| B | '3d element, 5th element, 6th element, 7th element, 8th element' | 'white, purple, teal, black' |
实现代码
1. 构造初始DataFrame
import pandas as pd data = { 'ID': ['A', 'A', 'B', 'B', 'B'], 'Elements': [ "'1st element, 2d element, 3d element'", "'2d element, 4th element'", "'3d element, 5th element, 6th element'", "'3d element, 5th element, 7th element'", "'3d element, 5th element, 8th element'" ], 'Colors': [ "'red, blue'", "'blue, green'", "'white, purple'", "'white, teal'", "'white, black'" ] } df = pd.DataFrame(data)
2. 定义字符串去重函数
这个函数负责处理单引号包裹的字符串,拆分元素、去重后重新拼接:
def deduplicate_string(s): # 移除首尾单引号 s_clean = s.strip("'") # 拆分元素并去除每个元素的前后空格 items = [item.strip() for item in s_clean.split(',')] # 去重并保留原始出现顺序 unique_items = list(dict.fromkeys(items)) # 重新拼接成带单引号的字符串 return f"'{', '.join(unique_items)}'"
3. 分组并应用去重逻辑
按ID分组,对Elements和Colors列分别应用去重函数:
result_df = df.groupby('ID').agg({ 'Elements': lambda x: deduplicate_string(', '.join(x)), 'Colors': lambda x: deduplicate_string(', '.join(x)) }).reset_index()
4. 查看最终结果
print(result_df)
运行后即可得到符合期望的DataFrame,其中dict.fromkeys的使用保证了元素的原始出现顺序,避免了用set去重导致的顺序混乱问题。
内容的提问来源于stack exchange,提问作者crocefisso
相关产品推荐
相关产品推荐

