You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

拼接Pandas数据框时如何保留MultiIndex中的Category数据类型?

保留MultiIndex中Categorical类型的拼接方案

问题背景

处理染色体名称这类需要自定义排序的数据时,通常会将其设为Categorical类型并纳入MultiIndex(排序逻辑:先按染色体,再按位置)。但使用pd.concat拼接DataFrame后,MultiIndex中的Categorical类型会丢失,导致排序不符合预期(例如示例中X8、X9被排在X10之后,而非预设的顺序)。

示例代码

import pandas as pd

df1 = pd.DataFrame({
    "A": pd.Categorical(
        ["X9", "X9", "X10", "X10"],
        categories=["X8", "X9", "X10"], ordered=True),
    "B": [1, 2, 1, 2],
    "C": ["9_1", "9_2", "10_1", "10_2"],
    "1": [1, 2, 3, 4]}
).set_index(["A", "B", "C"])
print(df1.index.dtypes)

df2 = pd.DataFrame({
    "A": pd.Categorical(
        ["X8", "X8", "X10", "X10"],
        categories=["X8", "X9", "X10"], ordered=True),
    "B": [1, 2, 1, 2],
    "C": ["8_1", "8_2", "10_1", "10_2"],
    "2": [1, 2, 3, 4]}
).set_index(["A", "B", "C"])
print(df2.index.dtypes)

df = pd.concat([df1, df2], axis=1).sort_index()
print(df.index.dtypes)
print(df.to_string())

示例输出

A    category
B       int64
C      object
dtype: object
A    category
B       int64
C      object
dtype: object
A    object
B     int64
C    object
dtype: object
              1    2
A   B C             
X10 1 10_1  3.0  3.0
    2 10_2  4.0  4.0
X8  1 8_1   NaN  1.0
    2 8_2   NaN  2.0
X9  1 9_1   1.0  NaN
    2 9_2   2.0  NaN

解决方案

方法1:拼接后重新转换索引层级为Categorical

拼接完成后,手动将MultiIndex的目标层级重新设置为Categorical类型,再执行排序:

import pandas as pd

# 定义统一的分类规则(可根据实际染色体名称调整)
chromosome_categories = ["X8", "X9", "X10"]

# 创建DataFrame
df1 = pd.DataFrame({
    "A": pd.Categorical(["X9", "X9", "X10", "X10"], categories=chromosome_categories, ordered=True),
    "B": [1, 2, 1, 2],
    "C": ["9_1", "9_2", "10_1", "10_2"],
    "1": [1, 2, 3, 4]}
).set_index(["A", "B", "C"])

df2 = pd.DataFrame({
    "A": pd.Categorical(["X8", "X8", "X10", "X10"], categories=chromosome_categories, ordered=True),
    "B": [1, 2, 1, 2],
    "C": ["8_1", "8_2", "10_1", "10_2"],
    "2": [1, 2, 3, 4]}
).set_index(["A", "B", "C"])

# 拼接DataFrame
df = pd.concat([df1, df2], axis=1)

# 重新设置MultiIndex的A层级为Categorical
new_index_levels = [
    pd.Categorical(df.index.get_level_values("A"), categories=chromosome_categories, ordered=True),
    df.index.get_level_values("B"),
    df.index.get_level_values("C")
]
df.index = pd.MultiIndex.from_arrays(new_index_levels, names=["A", "B", "C"])

# 按索引排序
df = df.sort_index()

print(df.index.dtypes)
print(df.to_string())

输出结果

A    category
B       int64
C      object
dtype: object
              1    2
A   B C             
X8  1 8_1   NaN  1.0
    2 8_2   NaN  2.0
X9  1 9_1   1.0  NaN
    2 9_2   2.0  NaN
X10 1 10_1  3.0  3.0
    2 10_2  4.0  4.0

方法2:使用join替代concat(适合索引重叠场景)

如果两个DataFrame的索引存在重叠,使用join可以更直接地保留索引的Categorical类型:

import pandas as pd

chromosome_categories = ["X8", "X9", "X10"]

df1 = pd.DataFrame({
    "A": pd.Categorical(["X9", "X9", "X10", "X10"], categories=chromosome_categories, ordered=True),
    "B": [1, 2, 1, 2],
    "C": ["9_1", "9_2", "10_1", "10_2"],
    "1": [1, 2, 3, 4]}
).set_index(["A", "B", "C"])

df2 = pd.DataFrame({
    "A": pd.Categorical(["X8", "X8", "X10", "X10"], categories=chromosome_categories, ordered=True),
    "B": [1, 2, 1, 2],
    "C": ["8_1", "8_2", "10_1", "10_2"],
    "2": [1, 2, 3, 4]}
).set_index(["A", "B", "C"])

# 使用join拼接,how="outer"保留所有索引
df = df1.join(df2, how="outer").sort_index()

print(df.index.dtypes)
print(df.to_string())

此方法的输出结果与方法1完全一致,且无需手动转换索引类型。

原理说明

pd.concat在处理MultiIndex时,若不同DataFrame的索引层级类型存在隐式转换(如Categorical与其他类型混合),会自动将Categorical降级为object类型。通过手动重置索引类型,或使用更注重索引兼容性的join方法,可避免该问题,确保自定义排序规则生效。


内容的提问来源于stack exchange,提问作者bli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 08:10:46