Python Pandas合并DataFrame行并跳过重叠行的实现方法
Got it, you want to stack your two DataFrames together but leave out any rows that show up in both. Here's a straightforward way to do this with pandas:
First, we’ll combine the two DataFrames vertically using pd.concat(), then remove any duplicate rows (these are the overlapping ones we want to skip).
Full Code Example
import numpy as np import pandas as pd # Your original data setup data1 = np.matrix([[4,75,2,5,84,2,6,5,554],[4,6,67,6,4,5,8,5,8]]).T data2 = np.matrix([[3,46,4,555,556,557,558,559,560],[1,2,4,1,3,5,3,1,5]]).T data1 = pd.DataFrame(data1) data2 = pd.DataFrame(data2) # Merge and remove overlapping rows combined_df = pd.concat([data1, data2]).drop_duplicates() print(combined_df)
How It Works
pd.concat([data1, data2])stacks the two DataFrames on top of each other, creating a single DataFrame with all rows from both sources..drop_duplicates()then scans for rows that are identical across all columns and removes the duplicates. By default, this keeps the first occurrence of each duplicate row (so if a row exists in both data1 and data2, the version from data1 stays in the final result).
Customization Option
If you want to exclude duplicate rows entirely (i.e., don’t keep any occurrence of rows that appear in both DataFrames), adjust the keep parameter:
combined_df = pd.concat([data1, data2]).drop_duplicates(keep=False)
In your specific sample data, there are no overlapping rows, so the combined result will include all 18 rows from both DataFrames. But this code will automatically handle any duplicates if they exist in your actual dataset.
内容的提问来源于stack exchange,提问作者Eric Kim

