You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python为层级数据生成指定统计指标?

层级CSV数据的统计与存储解决方案

针对你提出的需求,下面是一套实用的Python处理方案,同时解决CSV存储嵌套数据的痛点:

一、数据统计实现代码

先读取CSV数据,通过字典分组完成两个统计需求,代码简洁易维护:

import csv
from collections import defaultdict

# 1. 读取CSV原始数据
data = []
with open('location_data.csv', 'r', encoding='utf-8') as f:
    reader = csv.DictReader(f)
    for row in reader:
        data.append(row)

# 2. 统计每个国家关联的城市信息(自动去重)
country_city_stats = defaultdict(lambda: {'city_count': 0, 'cities': set()})
for row in data:
    country = row['country'].strip()
    city = row['city'].strip()
    country_city_stats[country]['cities'].add(city)
    country_city_stats[country]['city_count'] = len(country_city_stats[country]['cities'])
# 转成普通字典,把集合转成列表方便后续处理
country_city_stats = {k: {'city_count': v['city_count'], 'cities': list(v['cities'])} for k, v in country_city_stats.items()}

# 3. 统计每个城市下属的街道信息
city_address_stats = defaultdict(lambda: {'address_count': 0, 'addresses': []})
for row in data:
    city = row['city'].strip()
    address = row['address'].strip()
    city_address_stats[city]['addresses'].append(address)
    city_address_stats[city]['address_count'] = len(city_address_stats[city]['addresses'])
city_address_stats = dict(city_address_stats)

二、易使用的输出格式

1. 控制台友好输出

直接打印分组后的统计结果,清晰直观:

# 打印国家-城市统计
print("=== 国家关联城市统计 ===")
for country, stats in country_city_stats.items():
    print(f"国家:{country}")
    print(f"关联城市数量:{stats['city_count']}")
    print(f"具体城市:{', '.join(stats['cities'])}\n")

# 打印城市-街道统计
print("=== 城市下属街道统计 ===")
for city, stats in city_address_stats.items():
    print(f"城市:{city}")
    print(f"下属街道数量:{stats['address_count']}")
    print(f"具体街道:{', '.join(stats['addresses'])}\n")

输出效果示例:

=== 国家关联城市统计 ===
国家:usa
关联城市数量:2
具体城市:new york, seattle

国家:canada
关联城市数量:1
具体城市:ottawa

=== 城市下属街道统计 ===
城市:new york
下属街道数量:2
具体街道:100 Penn Plaza, 202 Barnes Ave

2. 存储方案优化

CSV是扁平格式,直接存嵌套结构确实麻烦,这里给两种实用方案:

方案一:拆分两个独立CSV表

把国家-城市、城市-街道的统计结果分别存成两个CSV,结构简单,后续读取和处理都方便:

# 保存国家-城市统计到CSV
with open('country_city_stats.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerow(['country', 'city_count', 'cities'])
    for country, stats in country_city_stats.items():
        writer.writerow([country, stats['city_count'], ', '.join(stats['cities'])])

# 保存城市-街道统计到CSV
with open('city_address_stats.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerow(['city', 'address_count', 'addresses'])
    for city, stats in city_address_stats.items():
        writer.writerow([city, stats['address_count'], ', '.join(stats['addresses'])])

生成的CSV每行对应一个统计项,多个城市/街道用逗号分隔,后续读取时可通过split(', ')还原成列表。

方案二:用JSON存储完整层级结构

如果需要保留父子关联的嵌套关系,JSON比CSV更适配,存储和读取都无需额外处理:

import json

# 整合两个统计结果
full_stats = {
    'country_city_stats': country_city_stats,
    'city_address_stats': city_address_stats
}

# 保存到JSON文件
with open('location_stats.json', 'w', encoding='utf-8') as f:
    json.dump(full_stats, f, indent=2, ensure_ascii=False)

生成的JSON文件可以直接用Python的json模块读取,拿到的就是完整的层级字典,非常适合后续程序调用。

总结

  • 统计需求通过字典分组即可轻松实现,defaultdict能简化初始化逻辑;
  • 若必须用CSV存储,拆分独立表是最实用的选择;
  • 要保留层级关系时,优先用JSON,天然适配嵌套数据结构。

内容的提问来源于stack exchange,提问作者Winner1235813213455

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:45:24