如何用Pandas实现纽约市树木统计及占比计算?
解决思路与代码优化
别用逐行遍历的方式,Pandas的核心优势就是向量化操作,既简洁又高效,完全能满足你的需求。以下是具体实现步骤:
1. 预计算全局与行政区的基础统计数据
先一次性算出所有必要的汇总数据,避免重复计算:
# 纽约市总树木数量 nyc_total_trees = len(Tree_data) # 各行政区的总树木数量(假设行政区列名为'borough',请替换为实际列名) borough_total_trees = Tree_data['borough'].value_counts() # 按树种+行政区分组的树木数量,缺失的组合填充0 tree_borough_counts = Tree_data.groupby(['spc_common', 'borough']).size().unstack(fill_value=0) # 各树种的全市总数 tree_nyc_counts = Tree_data['spc_common'].value_counts()
2. 重构你的formatInfo函数
直接用预计算好的数据生成统计结果,不用嵌套循环:
def formatInfo(matched_trees): for tree_name in matched_trees: print(f"Entry: {tree_name}") # 全市该树种总数与占比 total_nyc = tree_nyc_counts.get(tree_name, 0) pct_nyc = (total_nyc / nyc_total_trees) * 100 if nyc_total_trees > 0 else 0 print(f"Total in NYC: {total_nyc} ({pct_nyc:.2f}%)") # 各行政区的数量与占比 print("Borough breakdown:") for borough in borough_total_trees.index: total_borough = tree_borough_counts.loc[tree_name, borough] if tree_name in tree_borough_counts.index else 0 pct_borough = (total_borough / borough_total_trees[borough]) * 100 if borough_total_trees[borough] > 0 else 0 print(f" {borough}: {total_borough} ({pct_borough:.2f}%)") print("---")
为什么不用遍历?
- 逐行循环在数据量大的时候会非常慢,Pandas的内置方法都是底层优化过的,速度能提升几十甚至上百倍
- 分组统计、值计数这些都是Pandas原生功能,代码更简洁,逻辑也更清晰,后期维护更方便
如果你的行政区列名不是'borough',直接替换成数据集里的实际列名即可。测试时可以传入单个树种名列表,比如formatInfo(['red maple', 'pin oak'])。
内容的提问来源于stack exchange,提问作者Gabe Stewart-Guido
相关产品推荐
相关产品推荐

