You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何Pandas str.count比apply更慢?测试结果与认知不符

Pandas字符串统计性能疑问:str.count为何慢于apply?

我正在处理一个大型数据集,重点关注以下两列:

GenotypeIteration
100101100110111011010110000100111110111110000001110010011011111111011011110010110
000111000010110100000001100100101001011010110010101011101101110001010001100000000
001001001001001010001001011001101001011100001001110000000110010110011011110000110
100010101011000001011100010111110001101011001010101111001100110111010100111111100
110101010100100011101001101100010010101010011110001110111101101011010101000111100

我希望创建一个新列,统计Genotype列中包含多少个1。

尝试的两种方法

方法1:使用Pandas内置str模块

%%timeit
total_df['Count_1'] = total_df['Genotype'].str.count('1')

性能结果:10.9 s ± 183 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

方法2:使用apply()方法

%%timeit
total_df['Count_1'] = total_df['Genotype'].apply(lambda x: x.count('1'))

性能结果:2.63 s ± 13.8 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

疑问

第二种方法性能有显著提升,但根据认知,apply()方法通常比Pandas内置向量化方法更慢,我忽略了什么?

补充说明:使用的Pandas版本为pd.__version__ = 2.0.3


内容的提问来源于stack exchange,提问作者Tamames

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 20:20:24