如何将文本文件数据整理为嵌套字典以统计每日商品销量?
问题:如何将商品销量统计结果存入日期对应的嵌套字典?
原始数据格式
DAY:Monday Banana,Banana,Dragon Fruit Dragon Fruit,Ice Pops,Ice Pops Eggs,Dragon Fruit,Hamburger Buns,Dragon Fruit,Carrot,Apple,Banana Ice Pops,Carrot,Dragon Fruit,Banana,Eggs,Eggs,Eggs,Eggs DAY:Tuesday Banana,Hamburger Buns,Dragon Fruit,Ice Pops Hamburger Buns,Dragon Fruit,Dragon Fruit,Carrot,Apple,Carrot,Carrot,Eggs,Apple,Ice Pops Carrot,Ice Pops,Apple,Dragon Fruit,Ice Pops,Apple,Banana,Eggs Banana,Carrot,Eggs,Carrot,Eggs,Apple,Eggs Carrot,Eggs,Hamburger Buns,Dragon Fruit,Apple,Hamburger Buns,Carrot,Dragon Fruit Dragon Fruit,Ice Pops,Hamburger Buns,Hamburger Buns,Banana,Hamburger Buns,Carrot DAY:Wednesday Banana,Banana,Ice Pops Apple,Carrot,Hamburger Buns Apple,Carrot,Carrot,Carrot,Dragon Fruit,Carrot,Apple,Carrot,Dragon Fruit,Hamburger Buns Apple,Hamburger Buns,Dragon Fruit,Ice Pops
现有代码问题
你已分别实现了日期字典的创建和单行商品统计,但未将两者关联,导致统计结果无法对应到正确日期下。
创建日期主字典的代码
weekly_sales = {} source:str purchases = [] for line in data_purchases: if line.split(":")[0] == "DAY": #Had to separate between the header and the rest. Once this statement is done, it would remove the first row source = line.strip().rpartition(":")[-1] if source not in weekly_sales: weekly_sales[source] = {}
单行商品统计代码
for line in data_purchases: if line.split(":")[0] == "DAY": line = next(data_purchases) wordsCount = {} for item in line.split(",")[1:]: #.split so i can get each element in the lime item = item.strip() #strip \n from the elements. otherwise, the output would look like Dragon Fruit\n for instance if item not in wordsCount: wordsCount[item] = 1 else: wordsCount[item] += 1 else: #for every other line of the data wordsCount = {} for item in line.split(','): item = item.strip() if item not in wordsCount: wordsCount[item] = 1 else: wordsCount[item] += 1
解决方案
将两个逻辑合并为一个循环,通过current_day变量跟踪当前处理的日期,同时完成单行统计和结果关联,还能同时计算每行统计和每日总销量:
weekly_sales = {} current_day = None for line in data_purchases: line = line.strip() if not line: # 跳过空行 continue if line.startswith("DAY:"): # 提取当前日期名称 current_day = line.split(":")[-1].strip() # 初始化日期对应的嵌套结构:同时存储每行统计和每日总销量 weekly_sales[current_day] = { "daily_total": {}, # 每日各商品总销量 "line_stats": [] # 每行的商品统计结果列表 } else: if current_day is None: continue # 无对应日期的行直接跳过 # 统计当前行的商品出现次数 line_count = {} items = [item.strip() for item in line.split(",")] for item in items: # 用get方法简化计数逻辑,无需if-else判断 line_count[item] = line_count.get(item, 0) + 1 # 将单行统计结果存入当前日期的line_stats weekly_sales[current_day]["line_stats"].append(line_count) # 累计到每日总销量 for item, count in line_count.items(): weekly_sales[current_day]["daily_total"][item] = weekly_sales[current_day]["daily_total"].get(item, 0) + count
代码说明
- 跟踪当前日期:用
current_day变量记录正在处理的日期,避免分开循环导致的关联断层。 - 嵌套结构初始化:每个日期下同时维护
daily_total(累计总销量)和line_stats(每行明细统计),满足你两个核心需求。 - 简化计数逻辑:使用
dict.get(key, default)方法替代if-else判断,代码更简洁高效。 - 结果关联:处理非日期行时,直接将统计结果关联到当前日期对应的字典中,无需额外操作。
示例输出(部分)
打印weekly_sales["Monday"]["daily_total"]会得到:
{ 'Banana': 4, 'Dragon Fruit': 5, 'Ice Pops': 3, 'Eggs': 5, 'Hamburger Buns': 1, 'Carrot': 2, 'Apple': 1 }
内容的提问来源于stack exchange,提问作者Charl Margaux Elcano
相关产品推荐
相关产品推荐

