functools.partial相比lambda的性能优化具体是什么?
functools.partial is faster than lambda in this max() key function test Great question! I totally get why you’d assume partial is just a tool for binding arguments—let’s break down the performance optimizations that make it outperform the lambda in your test case.
First, let’s recap your test setup in CPython 3.6.4:
from functools import partial def add(x, y, z, a): return x + y + z + a list_of_as = list(range(10000)) def max1(): return max(list_of_as , key=lambda a: add(10, 20, 30, a)) def max2(): return max(list_of_as , key=partial(add, 10, 20, 30))
And your timing results confirm the clear performance gap:
In [2]: %timeit max1() 4.36 ms ± 42.3 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)
In [3]: %timeit max2() 3.67 ms ± 25.9 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)
Here’s why partial delivers better performance:
C-level vs. Python-level call overhead
functools.partialis implemented directly in C for CPython, while lambdas are full Python-level function objects. When you call the lambda, it has to execute its own Python bytecode to load theaddfunction, push the fixed arguments (10, 20, 30), add the passeda, then invokeadd. The partial object skips all this Python bytecode execution—its__call__method is handled in C, which directly forwards remaining arguments toaddwithout the extra Python function layer.No extra bytecode to run per call
Let’s look at the lambda’s underlying bytecode (using thedismodule):import dis lambda_key = lambda a: add(10,20,30,a) dis.dis(lambda_key)Output:
1 0 LOAD_GLOBAL 0 (add) 2 LOAD_CONST 1 (10) 4 LOAD_CONST 2 (20) 6 LOAD_CONST 3 (30) 8 LOAD_FAST 0 (a) 10 CALL_FUNCTION 4 12 RETURN_VALUEEvery time the lambda is called (10,000 times in your
max()call), all these bytecode steps run. The partial object doesn’t have this overhead—it’s pre-configured with the fixed arguments, so each call is a direct, optimized dispatch toadd.Efficient argument storage
When you create apartial, it stores the fixed arguments and target function in a structure optimized for quick retrieval during calls. The lambda, by contrast, has to re-fetch theaddreference and literal values (even though they’re cached) every time it runs, adding tiny but cumulative overhead.
In tight loops like the one in max(), these small per-call differences add up quickly—hence the ~16% speedup you see with partial. It’s not just a syntactic convenience; it’s an optimized tool for reducing function call overhead in performance-sensitive code.
内容的提问来源于stack exchange,提问作者pawelswiecki

