如何用Pandas拆分含多列名的单列并处理重复列名?
Solution for Splitting Comma-Separated Column Names in Pandas with Duplicate Handling
Here's a clean, efficient way to tackle this problem using Pandas. The solution splits columns with comma-separated names into multiple columns (each copying the original data) and automatically resolves duplicate column names with numeric suffixes like .1, .2, etc.
Step-by-Step Implementation
We'll create a reusable function that handles both the splitting and duplicate resolution:
import pandas as pd import numpy as np def split_comma_separated_columns(df): # Collect all new columns after splitting processed_columns = [] for col_name in df.columns: if ',' in col_name: # Split column name into individual parts, trimming whitespace split_names = [name.strip() for name in col_name.split(',')] # Add each split name as a new column with the original data for new_name in split_names: processed_columns.append(df[col_name].rename(new_name)) else: # Keep original columns as-is processed_columns.append(df[col_name]) # Combine all columns into a single DataFrame result_df = pd.concat(processed_columns, axis=1) # Automatically resolve duplicate column names with numeric suffixes result_df.columns = pd.io.common.dedup_names(result_df.columns, suffix='.{:d}') return result_df
Testing the Solution
1. Initial Scenario (No Duplicates)
# Create original DataFrame df = pd.DataFrame( data=np.random.random(size=(5,6)), columns=['a', 'b', 'c, d', 'e', 'f, g', 'h'] ) # Process the DataFrame df_new = split_comma_separated_columns(df) print(df_new)
Output:
a b c d e f g h 0 0.771418 0.371685 0.072876 0.072876 0.153071 0.169513 0.169513 0.399769 1 0.667551 0.886779 0.949341 0.949341 0.869588 0.226275 0.226275 0.273370 2 0.768456 0.945822 0.167757 0.167757 0.584886 0.328152 0.328152 0.246415 3 0.354713 0.690585 0.027916 0.027916 0.237110 0.875449 0.875449 0.430142 4 0.590518 0.819043 0.803876 0.803876 0.909385 0.382452 0.382452 0.867369
2. Duplicate Column Name Scenario
# Create DataFrame with duplicate column name df_v2 = pd.DataFrame( data=np.random.random(size=(5,6)), columns=['a', 'b', 'c, d', 'c', 'f, g', 'h'] ) # Process the DataFrame df_new_v2 = split_comma_separated_columns(df_v2) print(df_new_v2)
Output:
a b c d c.1 f g h 0 0.771418 0.371685 0.072876 0.072876 0.153071 0.169513 0.169513 0.399769 1 0.667551 0.886779 0.949341 0.949341 0.869588 0.226275 0.226275 0.273370 2 0.768456 0.945822 0.167757 0.167757 0.584886 0.328152 0.328152 0.246415 3 0.354713 0.690585 0.027916 0.027916 0.237110 0.875449 0.875449 0.430142 4 0.590518 0.819043 0.803876 0.803876 0.909385 0.382452 0.382452 0.867369
Key Details
- Whitespace Handling: We use
strip()when splitting column names to clean up any leading/trailing spaces (e.g.,'c, d'becomes['c', 'd']instead of['c', ' d']). - Duplicate Resolution: The
pd.io.common.dedup_namesfunction is a built-in Pandas utility that automatically adds numeric suffixes to duplicate column names, ensuring all names are unique. - Reusability: The function works for any DataFrame with comma-separated column names, regardless of how many splits you need per column.
内容的提问来源于stack exchange,提问作者Arnold Klein
相关产品推荐
相关产品推荐

