You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中处理大整数?避免指数存储与数值截断

问题分析与解决方案

你的问题核心是float64精度限制和Pandas列类型自动推断的行为:

  • 18位整数超过了float64的精确表示范围(float64仅能精确表示≤2^53≈9.007e15的整数),当列被自动推断为float64类型时,会丢失末尾精度,导致数值截断。
  • 循环中用loc动态添加列时,Pandas默认会将列类型设为float64,即使你赋值的是Python整数(Python int支持任意精度,但Pandas会自动转换为兼容的NumPy类型,此时超出float64精度的部分会丢失)。

解决方案1:提前初始化列类型为Int64(推荐,支持数值运算)

先创建指定类型的空列,避免Pandas自动推断为float64,确保整数精确存储:

import pandas as pd
import numpy as np

df1 = pd.DataFrame(data={'A':[123123123123123123, 234234234234234234, 345345345345345345], 'B':[11,22,33]})

# 提前初始化新列为Pandas可空整数类型Int64,支持精确存储np.int64范围内的整数
df1['new column'] = pd.Series(dtype='Int64')

for i in range(df1.shape[0]):
    df1.loc[i, 'new column'] = 222222222222222222

print(df1)

输出:

A   B       new column
0  123123123123123123  11  222222222222222222
1  234234234234234234  22  222222222222222222
2  345345345345345345  33  222222222222222222

解决方案2:用字符串类型存储(适合无需数值运算的场景)

如果不需要对该列进行数值计算,可以直接存储为字符串,完全保留原始格式:

import pandas as pd
import numpy as np

df1 = pd.DataFrame(data={'A':[123123123123123123, 234234234234234234, 345345345345345345], 'B':[11,22,33]})

# 初始化为字符串列
df1['new column'] = ''

for i in range(df1.shape[0]):
    df1.loc[i, 'new column'] = str(222222222222222222)

print(df1)

输出:

A   B       new column
0  123123123123123123  11  222222222222222222
1  234234234234234234  22  222222222222222222
2  345345345345345345  33  222222222222222222

补充说明

  • 之前的转换失败是因为:当列已经是float64类型时,精度已经丢失,再转成int64只能得到失真后的数值,无法恢复原始数据。必须在赋值前就指定正确的列类型。
  • 如果你的18位整数超过了np.int64的范围(>9223372036854775807),则需要用object类型存储Python原生整数(但不推荐频繁运算,会影响性能),或者用字符串。

内容的提问来源于stack exchange,提问作者naseefo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 06:45:53