使用Pandas将CSV转指定列字典,去除Color列重复值并排序
问题解决:处理CSV中Color列的重复值并排序生成字典
原始CSV数据
Flower;Price;Customer;Color Rose;20;Sam;Red, Yellow, Pink, Red Orchids;50;Jim;White, Purple, White Tulips;15;Genny;Red, White, Yellow, Yellow
需求
- 只保留
Flower和Color两列,丢弃其他列 - 清理
Color列:去掉重复颜色,按字母排序后存为列表 - 最终生成以花名为键、颜色列表为值的字典
当前已实现代码
import pandas as pd data = pd.read_csv(filename, usecols=['Flower','Color'], sep=';')
尝试过的无效代码
# 错误点:索引列名写错(应为Flower而非Name),参数格式不支持 data.set_index("Name").T.to_dict('sorted(set)') # 错误点:unique()和agg(set)的用法不符合当前数据结构的处理需求 data.unique().agg(set).to_dict()
期望输出
{ "Rose": ["Pink", "Red", "Yellow"], "Orchids": ["Purple", "White"], "Tulips": ["Red", "White", "Yellow"] }
正确实现代码及说明
你之前的代码没做「拆分颜色字符串、去重排序」这核心步骤,而且索引名也写错了。正确代码如下:
import pandas as pd data = pd.read_csv(filename, usecols=['Flower','Color'], sep=';') # 处理Color列:拆分字符串→去重→排序 data['Color'] = data['Color'].apply( lambda x: sorted(list(set(x.split(', ')))) ) # 转成目标字典格式 result = data.set_index('Flower')['Color'].to_dict() print(result)
代码解释
x.split(', '):把逗号加空格分隔的颜色字符串拆成单个颜色的列表set(...):自动去除列表里的重复颜色sorted(list(...)):把无序的集合转成有序列表,按字母排序set_index('Flower')['Color'].to_dict():将花名设为索引,直接转成键值对字典
内容的提问来源于stack exchange,提问作者prachi sutane
相关产品推荐
相关产品推荐

