为何Numpy会将object数组中的int类型转为float类型?
Numpy object数组与列表加法的异常类型转换问题
问题现象
当对dtype=object的Numpy数组与包含0的特定列表执行加法时,部分元素会从int转为float,导致精度丢失;但移除列表中的0或调整元素大小后,元素类型又恢复为int。
复现代码
含0列表的情况
X = np.array([5888275684537373439, 1945629710750298993],dtype=object) + [1158941147679947299,0] Y = np.array([5888275684537373439, 1945629710750298993],dtype=object) + [11589411476799472995,0] Z = np.array([5888275684537373439, 1945629710750298993],dtype=object) + [115894114767994729956,0] print(type(X[0]),X[0]) # <class 'int'> 7047216832217320738 print(type(Y[0]),Y[0]) # <class 'float'> 1.7477687161336848e+19 print(type(Z[0]),Z[0]) # <class 'int'> 121782390452532103395
移除0后的情况
X = np.array([5888275684537373439, 1945629710750298993],dtype=object) + [1158941147679947299] Y = np.array([5888275684537373439, 1945629710750298993],dtype=object) + [11589411476799472995] Z = np.array([5888275684537373439, 1945629710750298993],dtype=object) + [115894114767994729956] print(type(X[0]),X[0]) # <class 'int'> 7047216832217320738 print(type(Y[0]),Y[0]) # <class 'int'> 17477687161336846434 print(type(Z[0]),Z[0]) # <class 'int'> 121782390452532103395
更简洁的复现示例
import numpy as np A = np.array([1,1],dtype=object) + [2**62,0] B = np.array([1,1],dtype=object) + [2**63,0] C = np.array([1,1],dtype=object) + [2**64,0] D = np.array([1,1],dtype=object) + [2**63] E = np.array([1,1],dtype=object) + [2**63,2**63] print(type(A[0]),A[0]) # <class 'int'> 4611686018427387905 print(type(B[0]),B[0]) # <class 'float'> 9.223372036854776e+18 print(type(C[0]),C[0]) # <class 'int'> 18446744073709551617 print(type(D[0]),D[0]) # <class 'int'> 9223372036854775809 print(type(E[0]),E[0]) # <class 'int'> 9223372036854775809
原因分析
这不是Numpy的Bug,而是其类型提升机制导致的行为:
当Numpy执行数组与列表的加法时,会先将列表转换为Numpy数组。对于包含
0和大整数的列表,Numpy会优先尝试使用原生数值类型而非object类型:- 当列表中的大整数处于
int64范围外(比如2**63,超过有符号64位整数最大值),但同时存在0(默认int类型),Numpy会选择float64作为列表转数组后的类型——因为float64能近似容纳该大整数,而原生整数类型无法承载。 - 当列表移除
0或元素是更大的整数(比如2**64),Numpy找不到合适的原生数值类型容纳所有元素,就会将列表转为dtype=object的数组,此时加法操作会保留Python原生int类型,不会发生转换。
- 当列表中的大整数处于
当列表被转为
float64数组后,与原object数组执行加法时,Numpy会触发类型提升,将object数组的元素转为float64,导致大整数丢失精度;而如果列表被转为object数组,加法操作直接调用Python原生int的加法,保留精确值。验证:若显式将列表转为
dtype=object的数组再执行加法,就不会出现类型转换:B_fixed = np.array([1,1],dtype=object) + np.array([2**63,0], dtype=object) print(type(B_fixed[0]), B_fixed[0]) # <class 'int'> 9223372036854775809
内容的提问来源于stack exchange,提问作者Bobby Ocean
相关产品推荐
相关产品推荐

