嵌套字典实现按ID统计唯一产品数量
解决思路与实现方案
这思路完全靠谱!用双层字典来追踪每个ID对应的唯一产品,逻辑清晰又容易落地,我来帮你把这个想法细化成具体步骤和代码:
核心逻辑拆解
- 主字典:键为
DA列的ID,值是一个子字典(子字典的键为CB列的产品,值只需占位即可,比如True,我们只需要利用字典键的唯一性来自动去重) - 遍历所有数据行,逐个把ID和产品对应到字典结构里
- 最后通过子字典的长度,就能得到每个ID对应的唯一产品数量,再把这个数值填充到新的
DB列中
代码实现(以Python处理DataFrame为例)
方式一:双层字典(贴合你的原始思路)
import pandas as pd # 模拟你的数据,实际替换成你的数据集即可 df = pd.DataFrame({ 'DA': ['ID001', 'ID001', 'ID002', 'ID002', 'ID003', 'ID003', 'ID003'], 'CB': ['ProductX', 'ProductX', 'ProductY', 'ProductZ', 'ProductX', 'ProductY', 'ProductX'] }) # 初始化主字典 id_product_map = {} # 遍历每一行构建字典结构 for _, row in df.iterrows(): current_id = row['DA'] current_product = row['CB'] # 如果ID不在主字典里,新建一个空的子字典 if current_id not in id_product_map: id_product_map[current_id] = {} # 将产品作为子字典的键加入(自动去重) id_product_map[current_id][current_product] = True # 给DB列赋值:根据ID查对应子字典的长度 df['DB'] = df['DA'].apply(lambda x: len(id_product_map[x])) print(df)
方式二:用集合简化(更简洁,本质逻辑一致)
如果不想用双层字典,也可以用集合来存储每个ID对应的产品(集合天然去重),代码更简短:
import pandas as pd df = pd.DataFrame({ 'DA': ['ID001', 'ID001', 'ID002', 'ID002', 'ID003', 'ID003', 'ID003'], 'CB': ['ProductX', 'ProductX', 'ProductY', 'ProductZ', 'ProductX', 'ProductY', 'ProductX'] }) id_product_set = {} for _, row in df.iterrows(): current_id = row['DA'] current_product = row['CB'] if current_id not in id_product_set: id_product_set[current_id] = set() id_product_set[current_id].add(current_product) df['DB'] = df['DA'].map(lambda x: len(id_product_set[x])) print(df)
输出效果
运行后DB列会自动填充每个ID对应的唯一产品数:
| DA | CB | DB | |
|---|---|---|---|
| 0 | ID001 | ProductX | 1 |
| 1 | ID001 | ProductX | 1 |
| 2 | ID002 | ProductY | 2 |
| 3 | ID002 | ProductZ | 2 |
| 4 | ID003 | ProductX | 2 |
| 5 | ID003 | ProductY | 2 |
| 6 | ID003 | ProductX | 2 |
内容的提问来源于stack exchange,提问作者RugsKid
相关产品推荐
相关产品推荐

