You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效扁平化100×100规模的Pandas DataFrame?

Efficiently Flatten a Pandas DataFrame to Long Format

Hey Jack, great question! For turning your 100x100 square DataFrame into the long-format structure you described, the most efficient and clean approach uses Pandas' built-in stack() method—it’s optimized exactly for this kind of reshaping task, especially with larger datasets.

Step-by-Step Solution

First, let’s start with your sample DataFrame to demonstrate:

import pandas as pd
import numpy as np

# Your sample 3x3 DataFrame
df = pd.DataFrame(data=np.arange(1,10).reshape(3,3), index=['A', 'B', 'C'], columns=['A', 'B', 'C'])

To flatten it in one concise, efficient line:

flattened_df = df.stack().reset_index(name='val').rename(columns={'level_0': 'row', 'level_1': 'col'})

Breakdown of the Method

  • df.stack(): This collapses column labels into a multi-index paired with the original row index, creating a Series where each value is tied to its original row and column. For your 3x3 example, this gives a Series with a multi-index (row, col) and the corresponding values.
  • reset_index(name='val'): Converts the multi-index into regular columns (level_0 for rows, level_1 for columns) and names the value column val.
  • rename(...): Renames the auto-generated index columns to your desired row and col.

Result Verification

The output will match exactly what you’re looking for:

row col  val
0   A   A    1
1   A   B    2
2   A   C    3
3   B   A    4
4   B   B    5
5   B   C    6
6   C   A    7
7   C   B    8
8   C   C    9

Why This Is the Most Efficient

stack() leverages Pandas’ optimized internal vectorized operations (no slow Python-level loops) which makes it significantly faster than manual iteration or less specialized methods like melt() for square matrices. For your 100x100 dataset, this will generate 10,000 rows in milliseconds.

Alternative (Less Ideal for Square Matrices)

While melt() works, it requires resetting the row index first, adding an extra step:

flattened_melt = df.reset_index().melt(id_vars='index', var_name='col', value_name='val').rename(columns={'index': 'row'})

But stack() is more direct and efficient here since your row labels are already exactly the row values you need.

内容的提问来源于stack exchange,提问作者user6429576

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:33:17