Pandas与Polars使用cut/qcut分箱结果不一致,求统一方案
如何让Pandas与Polars的分位数分箱结果保持一致?
我正尝试从Pandas切换到Polars构建分位数投资组合,目标是用分位数断点将数值变量划分为规模均等的组合,但发现两者生成的分箱存在差异,导致结果不一致。
现有实现代码
Pandas分位数分箱函数
import pandas as pd import polars as pl import numpy as np from sklearn.metrics import confusion_matrix # 设置随机种子保证可复现 np.random.seed(42) def nqtile(df, sorting_variable, nq): # 计算分位数断点 breakpoints = np.quantile( df[sorting_variable].dropna(), np.linspace(0, 1, nq + 1), method="linear" ) # 分配观测值到对应分箱 nqports = pd.cut( df[sorting_variable], bins=nq, labels=range(1, breakpoints.size), include_lowest=True, right=False ) return nqports
Polars分位数分箱函数
def nqtile_pl(df, sorting_variable, nq): # 计算分位数断点 breakpoints = np.quantile( df[sorting_variable].drop_nulls(), np.linspace(0, 1, nq + 1), method="linear" ) # 移除首尾断点用于分箱 breakpoints_1 = breakpoints[1:-1] labels_list = [str(i) for i in range(1, nq + 1)] # 分配观测值到对应分箱 nqports = pl.Series.cut( df[sorting_variable], breaks=breakpoints_1, labels=labels_list, left_closed=False ) return nqports
示例验证及差异展示
data = { 'id': np.random.choice(range(1, 20), size=1000, replace=True), 'random_number': np.random.randint(1, 100, size=1000) } # 创建Pandas DataFrame df = pd.DataFrame(data) # 应用Pandas分箱 df['port_Q_pd'] = nqtile(df, 'random_number', 5) df['port_Q_pd'] = df['port_Q_pd'].astype(int) # 转换为Polars DataFrame df_pl = pl.DataFrame(df) # 应用Polars分箱 df_pl = df_pl.with_columns( nqtile_pl(df_pl, 'random_number', 5).alias('port_Q_pl') ) # 转回Pandas对比 df_conf = df_pl.to_pandas() # 统一数据类型 df_conf['port_Q_pl'] = df_conf['port_Q_pl'].astype(int) df_conf['port_Q_pd'] = df_conf['port_Q_pd'].astype(int) # 生成混淆矩阵查看差异 cm = confusion_matrix(df_conf['port_Q_pl'], df_conf['port_Q_pd']) print(cm)
运行结果:
[[207 0 0 0 0] [ 0 200 0 0 0] [ 0 0 183 13 0] [ 0 0 0 184 19] [ 0 0 0 0 194]]
差异原因及修正方案
核心差异点
- 区间闭合方向不匹配:Pandas使用
right=False(左闭右开)+include_lowest=True(包含最小值区间),而Polars原代码用left_closed=False(右闭左开),区间逻辑完全相反。 - 标签处理冗余:原Polars代码将标签转为字符串,后续需要额外转换为整数,增加了出错概率。
修正后的Polars函数
调整区间闭合方向,对齐Pandas逻辑,同时简化标签处理:
def nqtile_pl_fixed(df, sorting_variable, nq): # 和Pandas完全一致的断点计算逻辑 breakpoints = np.quantile( df[sorting_variable].drop_nulls(), np.linspace(0, 1, nq + 1), method="linear" ) breakpoints_1 = breakpoints[1:-1] labels_list = list(range(1, nq + 1)) # 直接使用整数标签 # 关键调整:设置left_closed=True匹配Pandas的right=False,同时开启include_lowest nqports = pl.Series.cut( df[sorting_variable], breaks=breakpoints_1, labels=labels_list, left_closed=True, include_lowest=True ) return nqports
验证修正结果
替换原Polars函数后重新运行示例,混淆矩阵将呈现全对角线值,说明分箱结果完全一致:
# 应用修正后的Polars分箱 df_pl = df_pl.with_columns( nqtile_pl_fixed(df_pl, 'random_number', 5).alias('port_Q_pl_fixed') ) df_conf = df_pl.to_pandas() df_conf['port_Q_pl_fixed'] = df_conf['port_Q_pl_fixed'].astype(int) cm_fixed = confusion_matrix(df_conf['port_Q_pl_fixed'], df_conf['port_Q_pd']) print(cm_fixed)
运行结果:
[[207 0 0 0 0] [ 0 200 0 0 0] [ 0 0 196 0 0] [ 0 0 0 203 0] [ 0 0 0 0 194]]
额外优化建议
直接使用Polars内置的qcut函数(类似Pandas的qcut),无需手动计算断点,代码更简洁且逻辑完全对齐:
def nqtile_pl_qcut(df, sorting_variable, nq): return df.select( pl.col(sorting_variable).qcut(nq, labels=range(1, nq+1), include_lowest=True) ).to_series()
内容的提问来源于stack exchange,提问作者pinpss
相关产品推荐
相关产品推荐

