如何避免Pandas保存与加载CSV文件时引入舍入误差?
解决CSV读写时的浮点数数值变化问题
你遇到的核心问题是:CSV作为文本格式,默认存储浮点数时会截断小数位,导致读取后与原数值产生微小差异。以下是几种可行的解决方法,可根据你的需求选择:
1. 直接指定高精度格式保存(最简单方案)
无需修改数据类型,只需在to_csv时通过float_format参数设置足够高的精度,确保能完整还原64位浮点数。%.17g是刚好能完整表示所有64位浮点数的格式:
import numpy as np import pandas as pd x = 10000 * np.pi df = pd.DataFrame({"test": [x]}) # 指定高精度格式保存CSV df.to_csv("pd_test.csv", float_format="%.17g") other_df = pd.read_csv("pd_test.csv") print(df["test"][0], other_df["test"][0]) print(df["test"][0] - other_df["test"][0])
运行后差值会变为0,实现读写数值完全一致。
2. 使用Decimal类型实现绝对精确存储
如果需要严格保证数值零误差,可以将数据转为decimal.Decimal类型,保存时以字符串形式写入CSV,读取后再转回Decimal:
import numpy as np import pandas as pd from decimal import Decimal # 先将浮点数值转为字符串再转Decimal,避免转换过程中的误差 x = Decimal(str(10000 * np.pi)) df = pd.DataFrame({"test": [x]}) # 直接以字符串形式保存 df.to_csv("pd_test.csv", index=False) # 读取时先以object类型加载,再转回Decimal other_df = pd.read_csv("pd_test.csv", dtype={"test": object}) other_df["test"] = other_df["test"].apply(Decimal) print(df["test"][0], other_df["test"][0]) print(df["test"][0] - other_df["test"][0])
这种方式能完全避免读写误差,但Decimal类型的计算效率低于numpy浮点数,适合对精度要求极高的场景。
3. 使用更高精度的浮点类型(兼容性需注意)
部分平台支持np.float128类型,其精度远高于默认的float64,配合高精度保存格式,能大幅降低读写误差:
import numpy as np import pandas as pd # 使用float128存储 x = np.float128(10000 * np.pi) df = pd.DataFrame({"test": [x]}) # 对应更高精度的保存格式 df.to_csv("pd_test.csv", float_format="%.20g") # 读取时指定float128类型 other_df = pd.read_csv("pd_test.csv", dtype={"test": np.float128}) print(df["test"][0], other_df["test"][0]) print(df["test"][0] - other_df["test"][0])
注意:float128并非所有环境都支持(比如Windows系统通常不支持),需根据运行环境判断是否可用。
内容的提问来源于stack exchange,提问作者Omroth
相关产品推荐
相关产品推荐

