为何拼接numpy数组比Python列表耗时更长?
str.join() drastically slower with a NumPy array than a Python list? I noticed this issue when working with random character sampling — let me know if my categorization is off. First, let's set up the test cases:
Generate a list of 1000 characters:
import random my_list = random.choices('abc', k=1000)
And a NumPy array from that list:
import numpy as np my_array = np.array(my_list)
When timing the operation to concatenate each into a single string, here's what I found:
''.join(my_list) # Vanilla Python list # 7.69 µs ± 103 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each) ''.join(my_array) # NumPy array # 257 µs ± 3.31 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
The question is: why is there such a massive difference in execution time?
My Initial Thoughts
At first, I guessed that str.join() might convert the NumPy array to a list under the hood before concatenating. But testing showed that's not the case — ndarray.tolist() is way too fast to account for the gap:
my_array.tolist() # 11.6 µs ± 87 ns per loop (mean ± std. dev. of 7 runs, 100000 loops each)
Could it be that str.join() is only highly optimized for Python lists, and not for NumPy arrays? That seems counterintuitive, since NumPy operations are usually more efficient than native Python implementations. I couldn't find the specific implementation of str.join() to dig deeper.
Additional Testing: Generating Random Characters with NumPy
For context, when generating random characters directly with NumPy:
my_array = np.random.choice([char for char in 'abc'], size=1000, replace=True) # 29.8 µs ± 2.33 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)
Comparing this to random.choices(), NumPy is an order of magnitude faster — other methods in the random module are even slower:
my_list = random.choices('abc', k=1000) # 276 µs ± 6.3 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)
This means when combining generation and concatenation, the total time for both approaches is roughly similar. But I'm confused why NumPy can't maintain its performance advantage when adding the concatenation step.
The Explanation
Great observation — let's unpack what's happening here:
How
str.join()works with Python lists:str.join()is heavily optimized for Python's built-in sequence types (like lists of strings). When you pass a list ofstrobjects, the method can directly access the underlying string data in bulk without any type conversion overhead. It pre-allocates the exact amount of memory needed for the final string and copies all characters in one go (or near one go), which is why it's so fast.The hidden overhead with NumPy arrays:
NumPy arrays are not Python sequences — they're homogeneous arrays of low-level C types. When you iterate over a NumPy array of characters (e.g., whenstr.join()loops through it), each element is converted from a NumPy scalar type (likenumpy.str_ornumpy.uint8, depending on your array's dtype) to a full Pythonstrobject. This conversion happens for every single element in the array, and that's where the massive overhead comes from.To put this in perspective: converting 1000 NumPy scalars to Python strings adds up to far more time than both generating the list and joining it. The
tolist()method is fast because it does this conversion in a more optimized bulk operation, butstr.join()can't take advantage of that — it has to do it element-by-element during iteration.How to fix the NumPy concatenation performance:
If you want to keep using NumPy and get fast concatenation, use NumPy's native string operations instead ofstr.join(). Thenp.char.cat()method is designed for this exact use case:np.char.cat(my_array) # ~1-2 µs per loop (way faster than both list join and array join!)This method operates directly on the NumPy array's underlying data without converting elements to Python objects, so it maintains NumPy's performance advantage.
Final Note on Total Workflow
When you combine generation and concatenation:
- For Python:
random.choices()+str.join()= ~276 + 7.7 = ~283.7 µs - For NumPy:
np.random.choice()+np.char.cat()= ~29.8 + 1 = ~30.8 µs
Now NumPy's performance advantage shines through — the key was using the right tool for the concatenation step instead of forcing str.join() to work with a NumPy array.
内容的提问来源于stack exchange,提问作者Engineero

