使用pd.concat合并DataFrame列报错:ValueError: cannot reindex from a duplicate axis
Hey there! Let's break down why you're hitting that error and get your column merge working right.
What's Causing the Error?
When you run df['E'] = pd.concat([df['C'], df['D']]), here's what's happening under the hood:
pd.concat([df['C'], df['D']])stacks the two columns vertically, creating a Series with 12 rows (6 values from column C, plus 6 from column D).- This stacked Series has a duplicate index: it repeats the original
0-5index twice (since both C and D inherit the index from your original df). - Your original DataFrame only has 6 rows. When you try to assign this 12-row Series to
df['E'], Pandas can't align the duplicate indices to your df's unique 0-5 index—hence thecannot reindex from a duplicate axiserror.
How to Fix It (Based on Your Goal)
Let's cover two common scenarios for what you might want to achieve:
Scenario 1: Stack C and D into a longer column (12 rows total)
If you want a new dataset where column E contains all of C's values followed by all of D's, don't assign to your original 6-row df. Instead, create a new DataFrame:
# Create a new df with just the stacked E column new_df = pd.DataFrame({'E': pd.concat([df['C'], df['D']]).reset_index(drop=True)}) # Or, if you want to keep the original columns repeated to match the 12 rows: new_df = pd.concat([df, df], ignore_index=True) new_df['E'] = pd.concat([df['C'], df['D']]).reset_index(drop=True)
The reset_index(drop=True) removes the duplicate original indices, so the new Series has a clean 0-11 index that aligns perfectly with the expanded df.
Scenario 2: Merge C and D values row-by-row into a single column
If you want each row in E to combine the C and D values from the same row (e.g., as a list, string, or calculated value), use row-wise operations instead:
# Merge into a list per row df['E'] = df.apply(lambda row: [row['C'], row['D']], axis=1) # Merge into a formatted string (convert numbers to strings first) df['E'] = df['C'].astype(str) + ", " + df['D'].astype(str) # Or calculate a combined value (e.g., max of C and D per row) df['E'] = df[['C', 'D']].max(axis=1)
Quick Note on Your Original Data
Looking at your sample df, the original E column matches C exactly. If that was intentional but you wanted to adjust it to include D, the row-wise methods above will help you tweak that to your needs.
内容的提问来源于stack exchange,提问作者Sarang Manjrekar

