如何在Pandas中处理大整数?避免指数存储与数值截断
问题分析与解决方案
你的问题核心是float64精度限制和Pandas列类型自动推断的行为:
- 18位整数超过了float64的精确表示范围(float64仅能精确表示≤2^53≈9.007e15的整数),当列被自动推断为float64类型时,会丢失末尾精度,导致数值截断。
- 循环中用
loc动态添加列时,Pandas默认会将列类型设为float64,即使你赋值的是Python整数(Python int支持任意精度,但Pandas会自动转换为兼容的NumPy类型,此时超出float64精度的部分会丢失)。
解决方案1:提前初始化列类型为Int64(推荐,支持数值运算)
先创建指定类型的空列,避免Pandas自动推断为float64,确保整数精确存储:
import pandas as pd import numpy as np df1 = pd.DataFrame(data={'A':[123123123123123123, 234234234234234234, 345345345345345345], 'B':[11,22,33]}) # 提前初始化新列为Pandas可空整数类型Int64,支持精确存储np.int64范围内的整数 df1['new column'] = pd.Series(dtype='Int64') for i in range(df1.shape[0]): df1.loc[i, 'new column'] = 222222222222222222 print(df1)
输出:
A B new column 0 123123123123123123 11 222222222222222222 1 234234234234234234 22 222222222222222222 2 345345345345345345 33 222222222222222222
解决方案2:用字符串类型存储(适合无需数值运算的场景)
如果不需要对该列进行数值计算,可以直接存储为字符串,完全保留原始格式:
import pandas as pd import numpy as np df1 = pd.DataFrame(data={'A':[123123123123123123, 234234234234234234, 345345345345345345], 'B':[11,22,33]}) # 初始化为字符串列 df1['new column'] = '' for i in range(df1.shape[0]): df1.loc[i, 'new column'] = str(222222222222222222) print(df1)
输出:
A B new column 0 123123123123123123 11 222222222222222222 1 234234234234234234 22 222222222222222222 2 345345345345345345 33 222222222222222222
补充说明
- 之前的转换失败是因为:当列已经是float64类型时,精度已经丢失,再转成int64只能得到失真后的数值,无法恢复原始数据。必须在赋值前就指定正确的列类型。
- 如果你的18位整数超过了np.int64的范围(>9223372036854775807),则需要用
object类型存储Python原生整数(但不推荐频繁运算,会影响性能),或者用字符串。
内容的提问来源于stack exchange,提问作者naseefo
相关产品推荐
相关产品推荐

