处理CSV分类平均值计算:替代元组字典的最优数据结构及更新方法
解决CSV分类平均尺寸计算中的数据结构问题
Hey there! Great question—totally get why you hit this roadblock with tuples being immutable. Let's walk through the best alternatives for your use case, along with practical code examples to implement them.
1. 字典 + 列表(最基础易上手)
因为列表是可变的(和元组正好相反),你可以轻松更新里面的总和和计数。我们可以把字典的每个值设为一个列表,格式为 [产品数量, 长度总和, 宽度总和, 高度总和, 重量总和]。
实现代码:
import csv # 初始化存储分类数据的字典 category_stats = {} with open('products.csv', 'r') as csv_file: # 用DictReader读取,方便按列名获取数据 reader = csv.DictReader(csv_file) for row in reader: category = row['Product Category'] # 把字符串类型的数值转为float(或int,根据你的数据类型) length = float(row['Length']) width = float(row['Width']) height = float(row['Height']) weight = float(row['Weight']) # 如果分类不在字典中,初始化列表 if category not in category_stats: category_stats[category] = [1, length, width, height, weight] else: # 已有分类,更新总和和计数 category_stats[category][0] += 1 category_stats[category][1] += length category_stats[category][2] += width category_stats[category][3] += height category_stats[category][4] += weight # 计算并输出平均尺寸 print("各分类平均尺寸:") for category, data in category_stats.items(): count, total_len, total_wid, total_hgt, total_wgt = data avg_len = total_len / count avg_wid = total_wid / count avg_hgt = total_hgt / count avg_wgt = total_wgt / count print(f"{category}: 平均长度={avg_len:.2f}, 平均宽度={avg_wid:.2f}, 平均高度={avg_hgt:.2f}, 平均重量={avg_wgt:.2f}")
2. 使用collections.defaultdict(更简洁)
defaultdict是Python标准库中的工具,它可以自动为不存在的键初始化默认值,省去了手动判断分类是否存在的步骤,让代码更简洁。
实现代码:
import csv from collections import defaultdict # 初始化defaultdict,默认值为[0, 0.0, 0.0, 0.0, 0.0](计数+各维度总和) category_stats = defaultdict(lambda: [0, 0.0, 0.0, 0.0, 0.0]) with open('products.csv', 'r') as csv_file: reader = csv.DictReader(csv_file) for row in reader: category = row['Product Category'] length = float(row['Length']) width = float(row['Width']) height = float(row['Height']) weight = float(row['Weight']) # 直接更新,无需判断分类是否存在 category_stats[category][0] += 1 category_stats[category][1] += length category_stats[category][2] += width category_stats[category][3] += height category_stats[category][4] += weight # 计算平均值的代码和第一种方法完全一致 print("各分类平均尺寸:") for category, data in category_stats.items(): count, total_len, total_wid, total_hgt, total_wgt = data avg_len = total_len / count avg_wid = total_wid / count avg_hgt = total_hgt / count avg_wgt = total_wgt / count print(f"{category}: 平均长度={avg_len:.2f}, 平均宽度={avg_wid:.2f}, 平均高度={avg_hgt:.2f}, 平均重量={avg_wgt:.2f}")
3. 使用dataclasses(可读性更强)
如果想让代码的可读性更高,变量名更清晰,可以用Python 3.7+引入的dataclass来定义一个统计类,把计数和各维度总和作为类的属性,这样代码更直观,不容易出错。
实现代码:
import csv from collections import defaultdict from dataclasses import dataclass # 定义统计数据的类 @dataclass class CategoryData: count: int = 0 total_length: float = 0.0 total_width: float = 0.0 total_height: float = 0.0 total_weight: float = 0.0 category_stats = defaultdict(CategoryData) with open('products.csv', 'r') as csv_file: reader = csv.DictReader(csv_file) for row in reader: category = row['Product Category'] stats = category_stats[category] stats.count += 1 stats.total_length += float(row['Length']) stats.total_width += float(row['Width']) stats.total_height += float(row['Height']) stats.total_weight += float(row['Weight']) # 计算平均值 print("各分类平均尺寸:") for category, stats in category_stats.items(): avg_len = stats.total_length / stats.count avg_wid = stats.total_width / stats.count avg_hgt = stats.total_height / stats.count avg_wgt = stats.total_weight / stats.count print(f"{category}: 平均长度={avg_len:.2f}, 平均宽度={avg_wid:.2f}, 平均高度={avg_hgt:.2f}, 平均重量={avg_wgt:.2f}")
总结
- 如果你是新手,先从字典+列表的方法入手,容易理解和调试;
- 想要更简洁的代码,用
defaultdict; - 追求代码可读性和可维护性,选
dataclasses。
这三种方法都能完美解决你之前遇到的元组不可变的问题,让你顺利更新分类的统计数据并计算平均值。
内容的提问来源于stack exchange,提问作者Alain
相关产品推荐
相关产品推荐

