使用正则表达式分组批量匹配多模式重命名Pandas DataFrame列
批量重命名DataFrame列名(多模式替换)
替换规则
male→m_female→f_working→wpopulation→popin the age group 0 to 6 years→_minor
方法一:使用pandas字符串方法链式调用
直接通过str.replace链式调用处理每个规则,按顺序执行可避免模式间干扰:
import pandas as pd # 初始化示例DataFrame cols_2 = ['state', 'population', 'male population', 'female population', 'working population', 'male working population', 'female working population', 'female population in the age group 0 to 6 years', 'male population in the age group 0 to 6 years', 'population in the age group 0 to 6 years'] df = pd.DataFrame(columns=cols_2) # 批量替换列名 df.columns = df.columns.str.replace('male ', 'm_')\ .str.replace('female ', 'f_')\ .str.replace('working ', 'w')\ .str.replace(' population', 'pop')\ .str.replace('in the age group 0 to 6 years', '_minor') # 查看结果 print(df.columns.tolist())
方法二:正则表达式+字典映射
如果需要更灵活的模式匹配逻辑,可结合re.sub与替换字典,遍历处理每个列名:
import pandas as pd import re cols_2 = ['state', 'population', 'male population', 'female population', 'working population', 'male working population', 'female working population', 'female population in the age group 0 to 6 years', 'male population in the age group 0 to 6 years', 'population in the age group 0 to 6 years'] df = pd.DataFrame(columns=cols_2) # 定义替换规则字典 replace_rules = { r'male ': 'm_', r'female ': 'f_', r'working ': 'w', r' population': 'pop', r'in the age group 0 to 6 years': '_minor' } # 定义单列名处理函数 def process_col_name(col): for pattern, replacement in replace_rules.items(): col = re.sub(pattern, replacement, col) return col # 应用到所有列名 df.columns = [process_col_name(col) for col in df.columns] # 查看结果 print(df.columns.tolist())
最终替换后的列名列表
['state', 'pop', 'm_pop', 'f_pop', 'wpop', 'm_wpop', 'f_wpop', 'f_pop_minor', 'm_pop_minor', 'pop_minor']
内容的提问来源于stack exchange,提问作者Adit Agrawal
相关产品推荐
相关产品推荐

