如何用Python为层级数据生成指定统计指标?
层级CSV数据的统计与存储解决方案
针对你提出的需求,下面是一套实用的Python处理方案,同时解决CSV存储嵌套数据的痛点:
一、数据统计实现代码
先读取CSV数据,通过字典分组完成两个统计需求,代码简洁易维护:
import csv from collections import defaultdict # 1. 读取CSV原始数据 data = [] with open('location_data.csv', 'r', encoding='utf-8') as f: reader = csv.DictReader(f) for row in reader: data.append(row) # 2. 统计每个国家关联的城市信息(自动去重) country_city_stats = defaultdict(lambda: {'city_count': 0, 'cities': set()}) for row in data: country = row['country'].strip() city = row['city'].strip() country_city_stats[country]['cities'].add(city) country_city_stats[country]['city_count'] = len(country_city_stats[country]['cities']) # 转成普通字典,把集合转成列表方便后续处理 country_city_stats = {k: {'city_count': v['city_count'], 'cities': list(v['cities'])} for k, v in country_city_stats.items()} # 3. 统计每个城市下属的街道信息 city_address_stats = defaultdict(lambda: {'address_count': 0, 'addresses': []}) for row in data: city = row['city'].strip() address = row['address'].strip() city_address_stats[city]['addresses'].append(address) city_address_stats[city]['address_count'] = len(city_address_stats[city]['addresses']) city_address_stats = dict(city_address_stats)
二、易使用的输出格式
1. 控制台友好输出
直接打印分组后的统计结果,清晰直观:
# 打印国家-城市统计 print("=== 国家关联城市统计 ===") for country, stats in country_city_stats.items(): print(f"国家:{country}") print(f"关联城市数量:{stats['city_count']}") print(f"具体城市:{', '.join(stats['cities'])}\n") # 打印城市-街道统计 print("=== 城市下属街道统计 ===") for city, stats in city_address_stats.items(): print(f"城市:{city}") print(f"下属街道数量:{stats['address_count']}") print(f"具体街道:{', '.join(stats['addresses'])}\n")
输出效果示例:
=== 国家关联城市统计 === 国家:usa 关联城市数量:2 具体城市:new york, seattle 国家:canada 关联城市数量:1 具体城市:ottawa === 城市下属街道统计 === 城市:new york 下属街道数量:2 具体街道:100 Penn Plaza, 202 Barnes Ave
2. 存储方案优化
CSV是扁平格式,直接存嵌套结构确实麻烦,这里给两种实用方案:
方案一:拆分两个独立CSV表
把国家-城市、城市-街道的统计结果分别存成两个CSV,结构简单,后续读取和处理都方便:
# 保存国家-城市统计到CSV with open('country_city_stats.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['country', 'city_count', 'cities']) for country, stats in country_city_stats.items(): writer.writerow([country, stats['city_count'], ', '.join(stats['cities'])]) # 保存城市-街道统计到CSV with open('city_address_stats.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['city', 'address_count', 'addresses']) for city, stats in city_address_stats.items(): writer.writerow([city, stats['address_count'], ', '.join(stats['addresses'])])
生成的CSV每行对应一个统计项,多个城市/街道用逗号分隔,后续读取时可通过split(', ')还原成列表。
方案二:用JSON存储完整层级结构
如果需要保留父子关联的嵌套关系,JSON比CSV更适配,存储和读取都无需额外处理:
import json # 整合两个统计结果 full_stats = { 'country_city_stats': country_city_stats, 'city_address_stats': city_address_stats } # 保存到JSON文件 with open('location_stats.json', 'w', encoding='utf-8') as f: json.dump(full_stats, f, indent=2, ensure_ascii=False)
生成的JSON文件可以直接用Python的json模块读取,拿到的就是完整的层级字典,非常适合后续程序调用。
总结
- 统计需求通过字典分组即可轻松实现,
defaultdict能简化初始化逻辑; - 若必须用CSV存储,拆分独立表是最实用的选择;
- 要保留层级关系时,优先用JSON,天然适配嵌套数据结构。
内容的提问来源于stack exchange,提问作者Winner1235813213455
相关产品推荐
相关产品推荐

