You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在指定列名列表中检测Pandas DataFrame的重复列

Detect Duplicate Columns in Pandas DataFrame Only for a Specified List of Columns

Great question! I've run into this exact scenario before—when you only care about duplicates in a specific subset of columns, the default df.columns.duplicated() just doesn't cut it because it checks every column. Let's fix this with a targeted approach that ignores duplicates outside your specified column list, even when Pandas auto-renames duplicates with suffixes like .1, .2, etc.

Step-by-Step Solution

The core idea is to:

  1. Extract the original column name from any auto-suffixed columns (e.g., turn col4.1 back into col4).
  2. Filter down to only the columns that belong to your target list.
  3. Check for duplicates among these filtered columns using their original names.

Full Code Implementation

import pandas as pd
import re

# 1. Load your CSV data into a DataFrame
df = pd.read_csv('your_data.csv')

# 2. Define your target column list (columns you want to check for duplicates)
target_columns = ['col1', 'col2', 'col3', 'col4']

# Helper function to reverse Pandas' auto-duplicate suffixing
def get_original_column_name(col_name):
    # Regex matches patterns like "col4.1" or "col4.2" and returns "col4"
    suffix_match = re.match(r'(.*?)\.\d+$', col_name)
    if suffix_match:
        return suffix_match.group(1)
    # Return original name if no numeric suffix exists
    return col_name

# 3. Filter columns to only those in your target list (including auto-suffixed duplicates)
relevant_columns = [
    col for col in df.columns 
    if get_original_column_name(col) in target_columns
]

# 4. Identify which original column names have duplicates
original_names = [get_original_column_name(col) for col in relevant_columns]
duplicated_original_names = {
    name for name in original_names 
    if original_names.count(name) > 1
}

# 5. Collect the actual duplicated column names (with suffixes)
duplicate_column_groups = []
for name in duplicated_original_names:
    group = [col for col in relevant_columns if get_original_column_name(col) == name]
    duplicate_column_groups.append(group)

# 6. Print results
print("Duplicate column groups found in target list:")
for group in duplicate_column_groups:
    base_name = get_original_column_name(group[0])
    print(f"- Original column '{base_name}' has duplicates: {group}")

How This Works

  • get_original_column_name: Uses regex to detect Pandas' auto-added numeric suffixes (like .1) and strips them to get back the original column name. It leaves columns without numeric suffixes untouched.
  • relevant_columns: Filters out any columns that aren't part of your target list—so duplicates of columns like custom (from your example) are completely ignored.
  • Duplicate Detection: We check which original column names appear more than once in the filtered list, then map those back to the actual column names (including suffixes) so you can see exactly which columns are duplicates.

Example Output

If your CSV has columns: col1, col2, col3, col4, col4.1, custom, custom, the output will be:

Duplicate column groups found in target list:
- Original column 'col4' has duplicates: ['col4', 'col4.1']

Notice that the duplicate custom columns are not included, since they're not in your target list.

内容的提问来源于stack exchange,提问作者blindside044

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 19:47:49