You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pd.concat合并DataFrame列报错:ValueError: cannot reindex from a duplicate axis

Fixing "ValueError: cannot reindex from a duplicate axis" When Merging Columns in Pandas

Hey there! Let's break down why you're hitting that error and get your column merge working right.

What's Causing the Error?

When you run df['E'] = pd.concat([df['C'], df['D']]), here's what's happening under the hood:

  • pd.concat([df['C'], df['D']]) stacks the two columns vertically, creating a Series with 12 rows (6 values from column C, plus 6 from column D).
  • This stacked Series has a duplicate index: it repeats the original 0-5 index twice (since both C and D inherit the index from your original df).
  • Your original DataFrame only has 6 rows. When you try to assign this 12-row Series to df['E'], Pandas can't align the duplicate indices to your df's unique 0-5 index—hence the cannot reindex from a duplicate axis error.

How to Fix It (Based on Your Goal)

Let's cover two common scenarios for what you might want to achieve:

Scenario 1: Stack C and D into a longer column (12 rows total)

If you want a new dataset where column E contains all of C's values followed by all of D's, don't assign to your original 6-row df. Instead, create a new DataFrame:

# Create a new df with just the stacked E column
new_df = pd.DataFrame({'E': pd.concat([df['C'], df['D']]).reset_index(drop=True)})

# Or, if you want to keep the original columns repeated to match the 12 rows:
new_df = pd.concat([df, df], ignore_index=True)
new_df['E'] = pd.concat([df['C'], df['D']]).reset_index(drop=True)

The reset_index(drop=True) removes the duplicate original indices, so the new Series has a clean 0-11 index that aligns perfectly with the expanded df.

Scenario 2: Merge C and D values row-by-row into a single column

If you want each row in E to combine the C and D values from the same row (e.g., as a list, string, or calculated value), use row-wise operations instead:

# Merge into a list per row
df['E'] = df.apply(lambda row: [row['C'], row['D']], axis=1)

# Merge into a formatted string (convert numbers to strings first)
df['E'] = df['C'].astype(str) + ", " + df['D'].astype(str)

# Or calculate a combined value (e.g., max of C and D per row)
df['E'] = df[['C', 'D']].max(axis=1)

Quick Note on Your Original Data

Looking at your sample df, the original E column matches C exactly. If that was intentional but you wanted to adjust it to include D, the row-wise methods above will help you tweak that to your needs.

内容的提问来源于stack exchange,提问作者Sarang Manjrekar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:18:48