You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效将含万余列的Pandas DataFrame所有列合并至第一列?

Efficiently Stack All Columns of a Large DataFrame into One Column

Hey there! Great question—when dealing with DataFrames that have tens of thousands of columns, the naive pd.concat approach can get pretty slow because it's repeatedly merging individual Series. Let's break down a much more efficient way to do this.

Why Your Original Method Is Slow

Your current approach:

b = pd.concat([a['col1'], ..., a['coln']]).reset_index(drop=True)

creates N separate Series and merges them one by one. Each merge involves copying data and adjusting the Series structure, which adds up exponentially when N is over 10,000. This leads to unnecessary overhead and slow performance.

The Optimal Solution: Leverage Numpy's Vectorized Operations

The fastest way to stack all columns vertically is to work directly with the underlying numpy array of your DataFrame. Here's the code:

import pandas as pd

# Example DataFrame (extend to N columns as needed)
cols = {'col1':['a','a','b','b'], 'coln':[1,2,3,4]}
a = pd.DataFrame(cols)

# Efficient column-wise stacking
b = pd.Series(a.values.ravel('F'), name='col1').reset_index(drop=True)

How This Works:

  • a.values: Grabs the raw numpy array from your DataFrame. This is a direct reference (no extra copy unless absolutely necessary).
  • ravel('F'): Flattens the array in column-major (Fortran) order—meaning we take all elements from the first column first, then the second, and so on. This exactly matches the stacking order you want.
  • Wrapping the flattened array in a pd.Series gives you the single-column structure, and reset_index(drop=True) ensures you get a clean sequential index.

Performance Comparison

For a DataFrame with 10,000 columns and 4 rows:

  • The numpy-based method runs in milliseconds.
  • The pd.concat approach would take seconds (or longer) due to repeated merge operations.

Notes on Data Types

If your columns have mixed data types, ravel('F') will automatically cast them to a compatible dtype (just like pd.concat does), so you don't have to worry about data loss or type mismatches.

Alternative (Pandas-Native) Method

If you prefer avoiding direct numpy access, you can use stack()—though it's slightly slower than the numpy approach for very large N:

b = a.stack().reset_index(drop=True).rename('col1')

This works by stacking columns into a hierarchical index, then flattening it. The overhead comes from handling the index structure, so it's less efficient for massive column counts.


内容的提问来源于stack exchange,提问作者Tomás Carrera de Souza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:47:47