为何bytes(lst)的执行速度显著慢于bytearray(lst)?
我用lst = [0] * 10**6做了基准测试,得到的执行时间如下:
5.4 ± 0.4 ms bytearray(lst) 5.6 ± 0.4 ms bytes(bytearray(lst)) 13.1 ± 0.7 ms bytes(lst)
使用的Python版本是:
3.13.0 (main, Nov 9 2024, 10:04:25) [GCC 14.2.1 20240910] namespace(name='cpython', cache_tag='cpython-313', version=sys.version_info(major=3, minor=13, micro=0, releaselevel='final', serial=0), hexversion=51183856, _multiarch='x86_64-linux-gnu')
原本我以为bytes(lst)和bytearray(lst)速度应该差不多,或者bytearray(lst)因为是可变类型功能更多,会更慢,但实际结果完全相反,甚至先转bytearray再转bytes的速度都比直接用bytes(lst)快很多。这到底是为什么呢?
背后的原因:CPython构造函数的实现差异
这个反直觉的结果,本质是CPython对bytearray和bytes的构造函数做了不同的优化:
bytearray(lst)的专属优化:CPython为bytearray处理整数列表的场景提供了专门的快速路径。它会直接分配一块连续的内存缓冲区,然后将列表中的整数批量拷贝到缓冲区中,全程几乎没有额外的冗余操作——校验和内存写入都是高效的批量处理,所以速度很快。bytes(lst)的通用路径开销:而bytes的构造函数在处理整数列表时,走的是通用的迭代器处理逻辑,没有针对整数列表做专门优化。它需要逐个遍历列表中的整数,做范围校验(虽然bytearray也会做,但bytes的校验流程更繁琐),并且因为bytes是不可变对象,在内存分配和数据固化的过程中,会多一些额外的步骤(比如可能需要先构建临时的可变缓冲区,再转换为不可变的bytes),这些额外操作累加起来就导致了明显的速度差异。bytes(bytearray(lst))的高效组合:这个路径之所以快,是因为它先借助bytearray(lst)的快速路径生成了连续的字节缓冲区,然后bytes()从bytearray转换时,只需要直接复制这块连续的内存即可——连续内存复制是CPU的强项,速度极快,所以两步操作的总耗时反而比直接用bytes(lst)处理零散整数列表要短。
基准测试脚本
from timeit import timeit from statistics import mean, stdev import random import sys setup = 'lst = [0] * 10**6' codes = [ 'bytes(lst)', 'bytearray(lst)', 'bytes(bytearray(lst))' ] times = {c: [] for c in codes} def stats(c): ts = [t * 1e3 for t in sorted(times[c])[:5]] return f'{mean(ts):5.1f} ± {stdev(ts):3.1f} ms ' for _ in range(25): random.shuffle(codes) for c in codes: t = timeit(c, setup, number=10) / 10 times[c].append(t) for c in sorted(codes, key=stats): print(stats(c), c) print('\nPython:') print(sys.version) print(sys.implementation)
备注:内容来源于stack exchange,提问作者Kelly Bundy

