You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas多索引列合并遇RuntimeWarning:int与tuple无法比较

多级列DataFrame合并的类型不兼容警告问题及解决

问题背景

有两个带pandas.MultiIndex(多级列)的DataFrame,二者索引无NaN/NaT但不完全一致,尝试以下三种列方向合并方式时均触发警告:

  • pd.concat([left_df, right_df], axis=1, sort=False)
  • pd.merge(left=left_df, right=right_df, left_index=True, right_index=True)
  • left_df.join(right_df)

警告内容:

RuntimeWarning: '<' not supported between instances of 'int' and 'tuple', sort order is undefined for incomparable objects.

警告原因

排查确认:左侧DataFrame的列元组第二个元素为字符串或字符串元组,右侧对应位置为整数;当两侧列元组元素类型完全一致时,无此警告。本质是不同类型的列索引元素无法进行大小比较,触发了Pandas内部排序逻辑的校验警告。

解决方案

通过统一两侧多级列索引的元素类型,消除类型不兼容问题,仅需1-2行代码即可完成:

# 将左侧列索引的所有元素转为字符串
left_df.columns = pd.MultiIndex.from_tuples([(str(l1), str(l2)) for l1, l2 in left_df.columns])
# 将右侧列索引的所有元素转为字符串
right_df.columns = pd.MultiIndex.from_tuples([(str(l1), str(l2)) for l1, l2 in right_df.columns])

执行上述转换后,再使用任意一种合并方式都不会触发警告。

是否为Pandas Bug?

这不属于Pandas未修复的Bug,而是预期行为。Pandas在合并过程中,即使设置sort=False,内部仍可能存在列索引的比较逻辑;当索引元素类型不可比较(如整数与元组)时,就会抛出该警告,属于Pandas对索引可比较性的正常校验。

可复现代码

import pandas as pd
import numpy as np

left_data = np.random.rand(1000, 5)
# 随机设置部分行/元素为NaN
for row in np.random.choice(np.arange(1000), size=10, replace=False):
    left_data[row, :] = np.nan
grid = np.array(np.meshgrid(np.arange(0, 1000), np.arange(0, 5))).T.reshape(-1, 2)
for coord in grid[np.random.choice(grid.shape[0], size=100, replace=False)]:
    left_data[coord] = np.nan
left_index = pd.date_range(start='20240101 09:00:00', periods=1000, freq='min')
left_columns = pd.MultiIndex.from_tuples([
    ('price', 'A'),
    ('price', 'B'),
    ('price', 'C'),
    ('diff', ('high', 'low')),
    ('diff', ('open', 'close'))
])
left_df = pd.DataFrame(data=left_data, index=left_index, columns=left_columns)

right_data = np.random.rand(990, 3)
# 随机设置部分行/元素为NaN
for row in np.random.choice(np.arange(990), size=5, replace=False):
    right_data[row, :] = np.nan
grid = np.array(np.meshgrid(np.arange(0, 990), np.arange(0, 3))).T.reshape(-1, 2)
for coord in grid[np.random.choice(grid.shape[0], size=50, replace=False)]:
    left_data[coord] = np.nan
right_index = pd.date_range(start='20240101 12:00:00', periods=990, freq='min')
right_columns = pd.MultiIndex.from_tuples([
    ('X', 1),
    ('X', 2),
    ('X', 3),
])
right_df = pd.DataFrame(data=right_data, columns=right_columns, index=right_index)

# 原合并代码(会触发警告)
df1 = pd.concat([left_df, right_df], axis=1, sort=False)
df2 = pd.merge(left_df, right_df, left_index=True, right_index=True)
df3 = left_df.join(right_df)

环境信息

  • Python版本:3.12.3
  • Pandas版本:2.2.3
  • NumPy版本:2.1.2
  • 操作系统:Ubuntu 24.04

内容的提问来源于stack exchange,提问作者kaddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 01:27:34