为何Pandas str.count比apply更慢?测试结果与认知不符
Pandas字符串统计性能疑问:
str.count为何慢于apply? 我正在处理一个大型数据集,重点关注以下两列:
| Genotype | Iteration |
|---|---|
| 10010110011011101101011000010011111011111000000111001001101111111101101111001011 | 0 |
| 00011100001011010000000110010010100101101011001010101110110111000101000110000000 | 0 |
| 00100100100100101000100101100110100101110000100111000000011001011001101111000011 | 0 |
| 10001010101100000101110001011111000110101100101010111100110011011101010011111110 | 0 |
| 11010101010010001110100110110001001010101001111000111011110110101101010100011110 | 0 |
我希望创建一个新列,统计Genotype列中包含多少个1。
尝试的两种方法
方法1:使用Pandas内置str模块
%%timeit total_df['Count_1'] = total_df['Genotype'].str.count('1')
性能结果:10.9 s ± 183 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
方法2:使用apply()方法
%%timeit total_df['Count_1'] = total_df['Genotype'].apply(lambda x: x.count('1'))
性能结果:2.63 s ± 13.8 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
疑问
第二种方法性能有显著提升,但根据认知,apply()方法通常比Pandas内置向量化方法更慢,我忽略了什么?
补充说明:使用的Pandas版本为pd.__version__ = 2.0.3
内容的提问来源于stack exchange,提问作者Tamames
相关产品推荐
相关产品推荐

