You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas concat函数的levels、keys、names参数作用及使用方法咨询

Understanding pd.concat(): Keys, Levels, and Practical Examples

Hey there! Let's break down Pandas' pd.concat()—the "Swiss Army knife" of data merging—and unpack those tricky parameters like keys and levels with plenty of hands-on examples to make everything click.

1. Basic pd.concat() Usage

First, let's cover the core functionality to set a foundation. pd.concat() lets you stack DataFrames/Series vertically (default, axis=0) or horizontally (axis=1). Here's a quick demo:

import pandas as pd

df1 = pd.DataFrame({'A': [1, 2], 'B': [3, 4]})
df2 = pd.DataFrame({'A': [5, 6], 'B': [7, 8]})

# Vertical stack (stack rows on top of each other)
vertical_stack = pd.concat([df1, df2])
print(vertical_stack)

# Horizontal stack (side-by-side columns)
horizontal_stack = pd.concat([df1, df2], axis=1)
print(horizontal_stack)

2. What does the keys parameter do?

The keys parameter adds a hierarchical MultiIndex to your concatenated result, so you can track exactly which original dataset each row/column came from. It's ideal for labeling groups when combining multiple related datasets.

Example 1: Keys for vertical concatenation

labeled_stack = pd.concat([df1, df2], keys=['Q1_Data', 'Q2_Data'])
print(labeled_stack)

The result will have a MultiIndex where the first level is your keys labels, and the second level is the original indexes of df1 and df2. You can slice by these labels easily:

# Pull all rows from Q1_Data
print(labeled_stack.loc['Q1_Data'])

Example 2: Keys for horizontal concatenation

When using axis=1, keys labels column groups instead of rows:

df3 = pd.DataFrame({'C': [9, 10], 'D': [11, 12]})
labeled_columns = pd.concat([df1, df3], axis=1, keys=['Original', 'New_Features'])
print(labeled_columns)

Now your columns will look like ('Original', 'A'), ('New_Features', 'C')—great for organizing related column sets.

3. What does the levels parameter do?

levels works with keys to define explicit values for each level of your MultiIndex, instead of letting Pandas infer them from your keys list. This is useful for enforcing consistency (like including labels for datasets you haven't even merged yet) or setting a specific order for index levels.

Example 1: Include missing labels in levels

Suppose you have three quarterly datasets, but only merge two—you can still include the third quarter in the index for future compatibility:

q1 = pd.DataFrame({'Sales': [100, 200]})
q2 = pd.DataFrame({'Sales': [300, 400]})
q3 = pd.DataFrame({'Sales': [500, 600]})  # Not merged, but we want it in the index

result = pd.concat(
    [q1, q2],
    keys=['Q1', 'Q2'],
    levels=[['Q1', 'Q2', 'Q3']],  # Define all possible quarter labels
    names=['Quarter']
)
print(result.index)

The index will now include Q3 even though we didn't merge that dataset—perfect for aligning with a predefined schema.

Example 2: Multi-level keys with custom levels

You can use levels for deeper MultiIndexes too, to enforce broader category labels:

usa = pd.DataFrame({'Revenue': [1000, 2000]})
canada = pd.DataFrame({'Revenue': [1500, 2500]})
uk = pd.DataFrame({'Revenue': [800, 1800]})

result = pd.concat(
    [usa, canada, uk],
    keys=[('North America', 'USA'), ('North America', 'Canada'), ('Europe', 'UK')],
    levels=[['North America', 'Europe'], ['USA', 'Canada', 'UK', 'France']],  # Include France even without data
    names=['Region', 'Country']
)
print(result.index)

4. Full Combined Parameter Examples

Let's put multiple parameters together for real-world scenarios:

Example: Combine keys, join, and ignore_index

df_left = pd.DataFrame({'A': [1, 2], 'B': [3, 4]})
df_right = pd.DataFrame({'B': [5, 6], 'C': [7, 8]})

# Merge vertically, label sources, keep only shared columns, preserve the MultiIndex
result = pd.concat(
    [df_left, df_right],
    keys=['Left_Side', 'Right_Side'],
    join='inner',  # Only keep columns present in both DataFrames
    ignore_index=False  # Keep the labeled MultiIndex
)
print(result)

Example: keys + levels + names for column indexes

result = pd.concat(
    [df1, df2],
    keys=['2022', '2023'],
    levels=[['2021', '2022', '2023']],  # Include 2021 for future data
    names=['Year'],
    axis=1
)
print(result)

内容的提问来源于stack exchange,提问作者piRSquared

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:07:31