You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

我的场景下遍历Pandas行完成列赋值的最优实现方式是什么?

Pandas行级计算效率优化方案

Pandas底层基于NumPy数组实现,支持向量化列运算,运算在C语言级别执行,完全不需要手动写Python级for循环遍历行,性能比你当前的实现高10~100倍,3万行数据可以做到秒级完成计算。

优化后代码

直接删除原有的for循环,替换为一行列运算即可:

belongs_node_df = pd.DataFrame.from_records(belongs_node, columns=['hashtag', 'tweets_id', 'tokenized_text','sentiment_compound'])
posted_node_df = pd.DataFrame.from_records(posted_node, columns=['username', 'num_followers', 'tweets_id'])
df_user_hashtag = pd.merge(posted_node_df, belongs_node_df, on='tweets_id', how='outer').sort_values('username')

# 替换原for循环的向量化运算
df_user_hashtag['p'] = 3 * df_user_hashtag['num_followers'] / df_user_hashtag['sentiment_compound']

注意事项

  • 原代码中的\为除号/的笔误,上述代码已做修正,若实际需求为其他运算可自行替换。
  • 建议提前处理sentiment_compound列的0值,避免计算得到无穷大inf,可补充异常值处理逻辑:
    import numpy as np
    # 示例:将无穷大、空值统一替换为0,可根据业务需求调整处理规则
    df_user_hashtag['p'] = df_user_hashtag['p'].replace([np.inf, -np.inf], 0).fillna(0)
    
  • 原代码的df_user_hashtag['p'][i]写法属于链式索引,容易触发SettingWithCopyWarning,存在修改失败的风险,不要在Pandas中使用这种赋值方式。
  • 如果后续有更复杂的行级逻辑无法通过简单算术运算实现,优先选择df.apply()方法,性能也远高于手动遍历行。

内容的提问来源于stack exchange,提问作者Minitorr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 09:27:03