You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何忽略索引横向拼接等长DataFrame?是否存在性能提升?

按行位置横向拼接DataFrame(忽略索引)的高效方法

Great question! I totally get what you're after—you want to stitch DataFrames side-by-side by their row positions, ignoring their original indexes entirely, and you're looking for a faster alternative to resetting indexes first (especially since you've seen how much faster numpy-based operations can be compared to pandas' index-aware logic).

Let's break this down:

Why pd.concat(axis=1) isn't working for you

By default, pd.concat(axis=1) aligns rows using their indexes. That's why you're getting those NaN values—df1 has indexes 0-9, df2 has 10-19, so there's no overlap, hence the missing values. And as you noted, ignore_index=True only resets the final result's index; it doesn't change the fact that pandas still aligns rows by their original indexes during the concat.

The numpy-based solution (fast & exactly what you need)

The trick here is to bypass pandas' index alignment entirely by working directly with the underlying numpy arrays. Here's how:

import pandas as pd
import numpy as np

df1 = pd.Series(range(10)).to_frame()
df2 = pd.Series(range(10), index=range(10, 20)).to_frame()

# Extract numpy arrays and stack horizontally, then convert back to DataFrame
result = pd.DataFrame(
    np.hstack([df1.to_numpy(), df2.to_numpy()]),
    columns=['df1_col', 'df2_col']  # Add your desired column names here
)

print(result)

This will give you a DataFrame where each row from df1 is paired with the corresponding row from df2, no NaNs, and it ignores both original indexes completely.

Performance boost: Why this is faster

Pandas' concat has to do a lot of heavy lifting: checking indexes, aligning rows, handling data type consistency across DataFrames, and more. By using np.hstack on the raw numpy arrays, you skip all that index-related overhead.

For small datasets the difference might be negligible, but as your DataFrames grow larger (think tens of thousands of rows or more), this numpy-based approach will be significantly faster—similar to the speedup you saw when using .values instead of pandas Series for arithmetic operations.

Quick note: .values vs .to_numpy()

While .values works, .to_numpy() is the preferred method in modern pandas (it's more explicit and handles dtypes more consistently). Either will get the job done, though.


内容的提问来源于stack exchange,提问作者The Unfun Cat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:34:09