如何高效将含万余列的Pandas DataFrame所有列合并至第一列?
Hey there! Great question—when dealing with DataFrames that have tens of thousands of columns, the naive pd.concat approach can get pretty slow because it's repeatedly merging individual Series. Let's break down a much more efficient way to do this.
Why Your Original Method Is Slow
Your current approach:
b = pd.concat([a['col1'], ..., a['coln']]).reset_index(drop=True)
creates N separate Series and merges them one by one. Each merge involves copying data and adjusting the Series structure, which adds up exponentially when N is over 10,000. This leads to unnecessary overhead and slow performance.
The Optimal Solution: Leverage Numpy's Vectorized Operations
The fastest way to stack all columns vertically is to work directly with the underlying numpy array of your DataFrame. Here's the code:
import pandas as pd # Example DataFrame (extend to N columns as needed) cols = {'col1':['a','a','b','b'], 'coln':[1,2,3,4]} a = pd.DataFrame(cols) # Efficient column-wise stacking b = pd.Series(a.values.ravel('F'), name='col1').reset_index(drop=True)
How This Works:
a.values: Grabs the raw numpy array from your DataFrame. This is a direct reference (no extra copy unless absolutely necessary).ravel('F'): Flattens the array in column-major (Fortran) order—meaning we take all elements from the first column first, then the second, and so on. This exactly matches the stacking order you want.- Wrapping the flattened array in a
pd.Seriesgives you the single-column structure, andreset_index(drop=True)ensures you get a clean sequential index.
Performance Comparison
For a DataFrame with 10,000 columns and 4 rows:
- The numpy-based method runs in milliseconds.
- The
pd.concatapproach would take seconds (or longer) due to repeated merge operations.
Notes on Data Types
If your columns have mixed data types, ravel('F') will automatically cast them to a compatible dtype (just like pd.concat does), so you don't have to worry about data loss or type mismatches.
Alternative (Pandas-Native) Method
If you prefer avoiding direct numpy access, you can use stack()—though it's slightly slower than the numpy approach for very large N:
b = a.stack().reset_index(drop=True).rename('col1')
This works by stacking columns into a hierarchical index, then flattening it. The overhead comes from handling the index structure, so it's less efficient for massive column counts.
内容的提问来源于stack exchange,提问作者Tomás Carrera de Souza

