更Pythonic的字符串清理方法:优化列名格式化实现
优化字符串清理为列名的Python方案
需求说明
将结构化字符串清理为规范列名,替代原有的多轮str.replace链式调用,实现更简洁、符合Python风格的写法。
示例输入
0 Estimate!!Total: 1 Estimate!!Total:!!Male: 2 Estimate!!Total:!!Male:!!Under 5 years 3 Estimate!!Total:!!Male:!!5 to 9 years
预期输出
0 Total: 1 Male: 2 Male_Under_5 3 Male_5_to_9 4 Male_10_to_14 5 Male_15_to_17
优化方案1:正则批量替换(简洁高效)
利用str.replace的正则模式,一次性整合多轮替换规则,大幅简化代码:
test['LABEL'] = test['LABEL'].str.replace( r'^(Estimate!!(Total:!!)?)|years|:!!|\s', lambda m: '_' if m.group() in (':!!', ' ') else '', regex=True ).str.rstrip('_')
规则解释
^(Estimate!!(Total:!!)?):匹配开头的Estimate!!或Estimate!!Total:!!,直接替换为空years:匹配并移除该文本:!!和\s:匹配后替换为下划线_str.rstrip('_'):移除末尾多余的下划线
优化方案2:拆分逻辑写法(可读性优先)
如果觉得单正则可读性差,可拆分规则分步处理,逻辑更清晰:
import re def clean_col_name(s): # 移除开头固定前缀 s = re.sub(r'^Estimate!!(Total:!!)?', '', s) # 替换分隔符和空格为下划线 s = re.sub(r':!!|\s', '_', s) # 移除years文本 s = s.replace('years', '') # 清理末尾下划线 return s.rstrip('_') test['LABEL'] = test['LABEL'].apply(clean_col_name)
内容的提问来源于stack exchange,提问作者Tinkinc
相关产品推荐
相关产品推荐

