You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何筛选数据集中指定多列值重复出现的行?

How to Filter Rows with Identical Values in Specified Columns

Absolutely! Filtering rows where specific feature columns have matching values is a common task, and it’s totally doable (even for your 10,000-row, 32-column dataset). Let’s use your example to walk through the best ways to handle this with Pandas (the go-to tool for tabular data in Python).

Step 1: Clarify the Objective

From your sample, you want to retain all rows where col2, col4, and col5 share identical values across multiple entries. In your test data, that means keeping rows 1,3,4 (where col2=2, col4=4, col5=5) and rows 2,5 (where col2=4, col4=6, col5=8).

Step 2: Practical Pandas Solutions

Here are two reliable methods to get your desired result:

Method 1: Use duplicated() (Simplest & Fastest)

The duplicated() method flags rows that are duplicates in your specified columns. Setting keep=False marks all duplicate rows (not just the first or last occurrence), which is exactly what we need for this task.

import pandas as pd

# Load your dataset (replace with pd.read_csv()/pd.read_excel() for your actual data)
# Using your sample data for demonstration:
sample_data = {
    'col1': [1, 3, 2, 4, 5, 2, 3],
    'col2': [2, 4, 2, 2, 4, 3, 4],
    'col3': [3, 3, 5, 7, 8, 1, 1],
    'col4': [4, 6, 4, 4, 6, 0, 5],
    'col5': [5, 8, 5, 5, 8, 9, 2]
}
df = pd.DataFrame(sample_data)

# Define the columns you want to match
target_cols = ['col2', 'col4', 'col5']

# Filter rows where the target columns have duplicate values
filtered_df = df[df.duplicated(subset=target_cols, keep=False)]

# Optional: Sort by target columns to group matching rows together for readability
filtered_df = filtered_df.sort_values(by=target_cols)

print(filtered_df)

This will output exactly the result you described:

col1 col2 col3 col4 col5
0 1 2 3 4 5
2 2 2 5 4 5
3 4 2 7 4 5
1 3 4 3 6 8
4 5 4 8 6 8

Method 2: Use groupby() (Flexible for Extra Processing)

If you want to perform additional operations on each group of matching rows (like calculating group stats), grouping by your target columns and filtering groups with multiple entries is a great approach:

# Group by the target columns
groups = df.groupby(target_cols)

# Keep only groups that have 2 or more rows
filtered_df = pd.concat([group for _, group in groups if len(group) >= 2])

# Optional sort to organize matching rows
filtered_df = filtered_df.sort_values(by=target_cols)

This gives the same end result as Method 1, but lets you easily extend logic if you need to work with each duplicate group individually.

For Your Large Dataset

Don’t stress about your 10k-row, 32-column dataset—Pandas handles this size effortlessly. Both methods will run in seconds, even with more columns. Just make sure to load your data correctly (e.g., pd.read_csv() for CSV files, pd.read_excel() for Excel spreadsheets).

内容的提问来源于stack exchange,提问作者sherek_66

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:32:16