如何在指定列名列表中检测Pandas DataFrame的重复列
Detect Duplicate Columns in Pandas DataFrame Only for a Specified List of Columns
Great question! I've run into this exact scenario before—when you only care about duplicates in a specific subset of columns, the default df.columns.duplicated() just doesn't cut it because it checks every column. Let's fix this with a targeted approach that ignores duplicates outside your specified column list, even when Pandas auto-renames duplicates with suffixes like .1, .2, etc.
Step-by-Step Solution
The core idea is to:
- Extract the original column name from any auto-suffixed columns (e.g., turn
col4.1back intocol4). - Filter down to only the columns that belong to your target list.
- Check for duplicates among these filtered columns using their original names.
Full Code Implementation
import pandas as pd import re # 1. Load your CSV data into a DataFrame df = pd.read_csv('your_data.csv') # 2. Define your target column list (columns you want to check for duplicates) target_columns = ['col1', 'col2', 'col3', 'col4'] # Helper function to reverse Pandas' auto-duplicate suffixing def get_original_column_name(col_name): # Regex matches patterns like "col4.1" or "col4.2" and returns "col4" suffix_match = re.match(r'(.*?)\.\d+$', col_name) if suffix_match: return suffix_match.group(1) # Return original name if no numeric suffix exists return col_name # 3. Filter columns to only those in your target list (including auto-suffixed duplicates) relevant_columns = [ col for col in df.columns if get_original_column_name(col) in target_columns ] # 4. Identify which original column names have duplicates original_names = [get_original_column_name(col) for col in relevant_columns] duplicated_original_names = { name for name in original_names if original_names.count(name) > 1 } # 5. Collect the actual duplicated column names (with suffixes) duplicate_column_groups = [] for name in duplicated_original_names: group = [col for col in relevant_columns if get_original_column_name(col) == name] duplicate_column_groups.append(group) # 6. Print results print("Duplicate column groups found in target list:") for group in duplicate_column_groups: base_name = get_original_column_name(group[0]) print(f"- Original column '{base_name}' has duplicates: {group}")
How This Works
get_original_column_name: Uses regex to detect Pandas' auto-added numeric suffixes (like.1) and strips them to get back the original column name. It leaves columns without numeric suffixes untouched.relevant_columns: Filters out any columns that aren't part of your target list—so duplicates of columns likecustom(from your example) are completely ignored.- Duplicate Detection: We check which original column names appear more than once in the filtered list, then map those back to the actual column names (including suffixes) so you can see exactly which columns are duplicates.
Example Output
If your CSV has columns: col1, col2, col3, col4, col4.1, custom, custom, the output will be:
Duplicate column groups found in target list: - Original column 'col4' has duplicates: ['col4', 'col4.1']
Notice that the duplicate custom columns are not included, since they're not in your target list.
内容的提问来源于stack exchange,提问作者blindside044
相关产品推荐
相关产品推荐

