np.vectorize为何会将一维数组中的np.uint8元素转换为int类型,却不改变列表元素的类型?
np.vectorize convert numpy array elements to Python scalars but not list elements? Great question! The behavior you're seeing boils down to how np.vectorize handles numpy array inputs vs. Python sequence inputs during its two key phases: the type-detecting test call, and the actual element-wise iteration. Let’s break this down step by step.
1. Recap the core behavior you observed
Your test function f prints the type of its input and checks for a .dtype attribute. When you pass a numpy array np.array([np.uint8(1)]) to np.vectorize(f):
- The first call to
fusesarr[0](anumpy.uint8scalar) to detect the output type → you see thenumpy.uint8type and its.dtype. - The second call uses an element from the array’s iterator, which gets converted to a Python
int→ hence the "no dtype exists" message.
But when you pass a list [np.uint8(1)] or a standalone numpy.uint8(1):
- Both the test call and the actual iteration use the original
numpy.uint8scalar → you see thenumpy.uint8type twice.
2. The root cause: How np.vectorize iterates over inputs
For numpy arrays:
np.vectorize uses numpy’s built-in iterator (np.nditer) to loop over array elements. By default, this iterator converts numpy scalar types (like np.uint8) to their equivalent Python native types (like int). This is a design choice to make numpy arrays feel more "Pythonic" when iterated over in standard loops.
You can verify this behavior outside of np.vectorize:
arr = np.array([np.uint8(1)]) for elem in arr: print(type(elem)) # Output: <class 'int'>
For Python lists (or other non-numpy sequences):
np.vectorize simply iterates over the list using standard Python iteration, which returns exactly the elements stored in the list—no type conversion happens. So the np.uint8 scalar you put into the list stays a np.uint8 scalar when pulled out.
Again, verify this outside np.vectorize:
lst = [np.uint8(1)] for elem in lst: print(type(elem)) # Output: <class 'numpy.uint8'>
3. The test call vs. actual iteration distinction
The test call (to determine output dtype/shape) always uses direct indexing (e.g., arr[0] for arrays, lst[0] for lists). Direct indexing on a numpy array returns a numpy scalar, which is why your first print shows numpy.uint8. It’s only during the actual iteration phase that the numpy iterator converts the scalar to a Python type.
4. A quick note on np.vectorize limitations
Keep in mind that np.vectorize is mostly a convenience function—it doesn’t actually speed up your code (it’s still a Python loop under the hood). If you need true vectorized operations that preserve numpy types, you’re better off using numpy’s native vectorized functions or writing your own ufuncs.
内容的提问来源于stack exchange,提问作者zabop

