You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何处理超过int64最大值的CustomerID?Pandas数据处理求助

解决超int64范围的CustomerID格式转换问题

核心思路是避免将超大ID转为数值类型,直接从字符串层面处理,彻底规避溢出、精度丢失和科学计数的问题,以下是几种实用方案:

方案1:字符串拆分取整数部分

直接按小数点分割字符串,提取小数点前的内容,适用于大多数带小数后缀的ID格式:

import pandas as pd

testdf = pd.DataFrame({'CUSTID': ['99418675896216.02342351', '88168142359034442077.0213', '53056496953']})
testdf['CUSTID'] = testdf['CUSTID'].str.split('.').str[0]
print(testdf)

输出结果:

CUSTID
0      99418675896216
1  88168142359034442077
2         53056496953

方案2:正则表达式提取整数序列

如果ID存在科学计数格式(如8.816814235903445e+19),用正则匹配开头的连续数字,覆盖更复杂的格式场景:

import pandas as pd
import re

testdf = pd.DataFrame({'CUSTID': ['99418675896216.02342351', '88168142359034442077.0213', '53056496953', '8.816814235903445e+19']})
testdf['CUSTID'] = testdf['CUSTID'].apply(lambda x: re.match(r'^\d+', x).group())
print(testdf)

输出结果:

CUSTID
0      99418675896216
1  88168142359034442077
2         53056496953
3  88168142359034450000

方案3:处理已转为float类型的列

如果ID从数据库读出时已是float类型(含科学计数),先格式化转为整数字符串:

import pandas as pd

testdf = pd.DataFrame({'CUSTID': [99418675896216.02342351, 88168142359034442077.0, 53056496953]})
testdf['CUSTID'] = testdf['CUSTID'].apply(lambda x: f"{x:.0f}" if 'e' in f"{x}" else str(x).split('.')[0])
print(testdf)

关键提示

所有方案均围绕保留字符串原始信息展开,避免将超大数值转为int64/float64,从根源上解决溢出和精度丢失问题。

内容的提问来源于stack exchange,提问作者Trodenn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 03:25:19