Python中round处理的变量两次打印结果不一致的原因及解决
Python中round处理后变量打印格式不一致的问题
代码示例
import xgboost as xgb from sklearn.model_selection import train_test_split import pandas as pd import numpy as np # Obtaining the data set, method 1 from sklearn.datasets import load_boston boston = load_boston() X = boston.data y = boston.target # Obtaining the data set, method 2 # data_url = "http://lib.stat.cmu.edu/datasets/boston" # raw_df = pd.read_csv(data_url, sep="\s+", skiprows=22, header=None) # X = np.hstack([raw_df.values[::2, :], raw_df.values[1::2, :2]]) # y = raw_df.values[1::2, 2] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) dtrain = xgb.DMatrix(X_train, label=y_train) dtest = xgb.DMatrix(X_test, label=y_test) model = xgb.train({'objective': 'reg:squarederror'}, dtrain) y_pred = model.predict(dtest) mean_col1 = round(y_test.mean(), 4) mean_col2 = round(y_pred.mean(), 4) # first print print(mean_col1, mean_col2) # second print print(f"real price avg: {mean_col1}, predict price avg: {mean_col2}")
运行输出
21.4882 20.5224 real price avg: 21.4882, predict price avg: 20.52239990234375
问题
为何第二次打印时mean_col2未保留4位小数?同一变量首次打印显示4位小数,二次打印却结果不同,该问题在Jupyter Notebook、IPython和py文件中均存在,且仅mean_col2出现此情况?
原因分析
- 数据类型差异:
y_test是numpy.ndarray类型,其均值返回numpy.float64,round后仍为该类型;而y_pred是XGBoost返回的numpy.float32类型,均值及round后结果也保持numpy.float32类型。 - 打印机制区别:
- 直接用
print输出多个变量时,numpy类型会调用自身的默认字符串表示逻辑,numpy.float32会自动截断到易读的小数位数(这里显示4位)。 - f-string中会将numpy类型转换为Python原生
float,但float32精度有限,20.5224无法被精确存储,实际存储的是近似值20.52239990234375,f-string会输出这个精确的存储值。
- 直接用
mean_col1是numpy.float64类型,21.4882可以被该类型精确表示,所以两种打印方式结果一致。
解决方法
有三种可行的处理方式:
- 转为Python原生float:在
round后将mean_col2转换为原生float,消除numpy类型的影响:mean_col2 = float(round(y_pred.mean(), 4)) - f-string指定小数位数:直接在格式化时控制显示精度,无需依赖
round结果:print(f"real price avg: {mean_col1:.4f}, predict price avg: {mean_col2:.4f}") - 用.item()提取原生值:通过numpy标量的
.item()方法获取Python原生数值后再round:mean_col2 = round(y_pred.mean().item(), 4)
内容的提问来源于stack exchange,提问作者iris_hu
相关产品推荐
相关产品推荐

