Python实现按contract统计唯一custId数量(不使用Pandas)
解决按contract统计唯一custId数量的问题(无Pandas版)
Hey there! Since you're a Python newbie and want to stick with the csv module instead of Pandas, let's build this functionality step by step. Here's a straightforward approach that uses basic Python data structures to get the job done:
完整实现代码
import csv # 初始化一个字典:键为contract编号,值为存储唯一custId的集合(集合自动去重) contract_unique_custs = {} with open('data.csv', 'r') as f: reader = csv.reader(f) # 跳过表头行(第一行的列名) next(reader) for row in reader: # 从每行数据中提取custId和contract(注意列的索引:custId是第0列,contract是第1列) cust_id = row[0] contract_num = row[1] # 如果当前contract还没在字典里,先创建一个空集合 if contract_num not in contract_unique_custs: contract_unique_custs[contract_num] = set() # 将custId加入对应contract的集合(集合会自动忽略重复的id) contract_unique_custs[contract_num].add(cust_id) # 输出统计结果 print("=== 按contract统计的唯一custId数量 ===") for contract, cust_ids in contract_unique_custs.items(): print(f"Contract {contract}: {len(cust_ids)} 个唯一用户")
代码逻辑拆解
- 数据结构选择:用字典
contract_unique_custs来关联每个contract和它对应的用户集合,集合的特性刚好能帮我们自动去重,不用手动判断重复的custId。 - 跳过表头:用
next(reader)读取并跳过第一行的列名,避免把表头当成数据处理。 - 遍历处理每行:从每行中提取需要的两个字段,更新字典里的集合。
- 统计输出:最后遍历字典,用
len(cust_ids)得到每个contract对应的唯一用户数量。
运行结果示例
=== 按contract统计的唯一custId数量 ===
Contract 2345: 2 个唯一用户
Contract 2346: 2 个唯一用户
内容的提问来源于stack exchange,提问作者JustCurious
相关产品推荐
相关产品推荐

