如何忽略索引横向拼接等长DataFrame?是否存在性能提升?
Great question! I totally get what you're after—you want to stitch DataFrames side-by-side by their row positions, ignoring their original indexes entirely, and you're looking for a faster alternative to resetting indexes first (especially since you've seen how much faster numpy-based operations can be compared to pandas' index-aware logic).
Let's break this down:
Why pd.concat(axis=1) isn't working for you
By default, pd.concat(axis=1) aligns rows using their indexes. That's why you're getting those NaN values—df1 has indexes 0-9, df2 has 10-19, so there's no overlap, hence the missing values. And as you noted, ignore_index=True only resets the final result's index; it doesn't change the fact that pandas still aligns rows by their original indexes during the concat.
The numpy-based solution (fast & exactly what you need)
The trick here is to bypass pandas' index alignment entirely by working directly with the underlying numpy arrays. Here's how:
import pandas as pd import numpy as np df1 = pd.Series(range(10)).to_frame() df2 = pd.Series(range(10), index=range(10, 20)).to_frame() # Extract numpy arrays and stack horizontally, then convert back to DataFrame result = pd.DataFrame( np.hstack([df1.to_numpy(), df2.to_numpy()]), columns=['df1_col', 'df2_col'] # Add your desired column names here ) print(result)
This will give you a DataFrame where each row from df1 is paired with the corresponding row from df2, no NaNs, and it ignores both original indexes completely.
Performance boost: Why this is faster
Pandas' concat has to do a lot of heavy lifting: checking indexes, aligning rows, handling data type consistency across DataFrames, and more. By using np.hstack on the raw numpy arrays, you skip all that index-related overhead.
For small datasets the difference might be negligible, but as your DataFrames grow larger (think tens of thousands of rows or more), this numpy-based approach will be significantly faster—similar to the speedup you saw when using .values instead of pandas Series for arithmetic operations.
Quick note: .values vs .to_numpy()
While .values works, .to_numpy() is the preferred method in modern pandas (it's more explicit and handles dtypes more consistently). Either will get the job done, though.
内容的提问来源于stack exchange,提问作者The Unfun Cat

