You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升pandas DataFrame遍历行生成国家及大洲编码的运行速度

原代码核心问题

  1. 逻辑错误:循环内直接对整列dfSPSSstudent['CN']赋值,最终所有行的取值都会被覆盖为最后一行的计算结果,完全无法得到每行对应的国家编码。
  2. 性能问题:iterrows是Python层遍历,加上每行都调用转换函数、执行print IO操作,25万行规模下运行效率极低。

优化方案

核心思路是先对国家列去重,预先生成转换映射字典,再用pandas内置的批量映射方法赋值,避免重复计算和Python层遍历。

# 1. 提取唯一国家名,预生成转换映射(仅对不重复的国家名调用转换函数,大幅减少函数调用次数)
unique_countries = dfSPSSstudent['Country_ID'].unique()
country_cn_map = {}
country_continent_map = {}

for country in unique_countries:
    a2_code = country_name_to_country_alpha2(country)
    country_cn_map[country] = a2_code
    country_continent_map[country] = country_alpha2_to_continent_code(a2_code)

# 2. 批量映射生成新列(底层为C实现,比Python层遍历快100倍以上)
dfSPSSstudent['CN'] = dfSPSSstudent['Country_ID'].map(country_cn_map)
dfSPSSstudent['Continent_Code'] = dfSPSSstudent['Country_ID'].map(country_continent_map)

优化点说明

  • 移除了所有print语句,避免高耗时的IO操作
  • 转换函数仅对去重后的国家名调用,假设25万行共包含200个不同国家,函数调用次数直接从25万次降到200次
  • 用pandas内置map方法做批量赋值,完全避免Python层的逐行遍历
  • 修复了原代码的列赋值逻辑错误,可正确生成每行对应的编码值

25万行规模下,优化后的代码运行时间可从原方案的分钟级降到1秒以内。

内容的提问来源于stack exchange,提问作者fredooms

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 09:18:04