如何处理超过int64最大值的CustomerID?Pandas数据处理求助
解决超int64范围的CustomerID格式转换问题
核心思路是避免将超大ID转为数值类型,直接从字符串层面处理,彻底规避溢出、精度丢失和科学计数的问题,以下是几种实用方案:
方案1:字符串拆分取整数部分
直接按小数点分割字符串,提取小数点前的内容,适用于大多数带小数后缀的ID格式:
import pandas as pd testdf = pd.DataFrame({'CUSTID': ['99418675896216.02342351', '88168142359034442077.0213', '53056496953']}) testdf['CUSTID'] = testdf['CUSTID'].str.split('.').str[0] print(testdf)
输出结果:
CUSTID 0 99418675896216 1 88168142359034442077 2 53056496953
方案2:正则表达式提取整数序列
如果ID存在科学计数格式(如8.816814235903445e+19),用正则匹配开头的连续数字,覆盖更复杂的格式场景:
import pandas as pd import re testdf = pd.DataFrame({'CUSTID': ['99418675896216.02342351', '88168142359034442077.0213', '53056496953', '8.816814235903445e+19']}) testdf['CUSTID'] = testdf['CUSTID'].apply(lambda x: re.match(r'^\d+', x).group()) print(testdf)
输出结果:
CUSTID 0 99418675896216 1 88168142359034442077 2 53056496953 3 88168142359034450000
方案3:处理已转为float类型的列
如果ID从数据库读出时已是float类型(含科学计数),先格式化转为整数字符串:
import pandas as pd testdf = pd.DataFrame({'CUSTID': [99418675896216.02342351, 88168142359034442077.0, 53056496953]}) testdf['CUSTID'] = testdf['CUSTID'].apply(lambda x: f"{x:.0f}" if 'e' in f"{x}" else str(x).split('.')[0]) print(testdf)
关键提示
所有方案均围绕保留字符串原始信息展开,避免将超大数值转为int64/float64,从根源上解决溢出和精度丢失问题。
内容的提问来源于stack exchange,提问作者Trodenn
相关产品推荐
相关产品推荐

