Python:修复DataFrame中提取并替换月份缩写的代码问题
问题分析与解决
原代码的问题
你的代码逻辑完全写反了,导致条件永远不成立,所以DataFrame没有变化:
any(abv in month_abv for abv in test.loc[inx,'month'])这段是把文件名拆成单个字符逐个检查是否在月份缩写列表里,单个字符不可能匹配到3个字母的月份缩写,所以条件永远为False,不会执行赋值操作。- 就算条件成立,
abv最后是文件名的最后一个字符,不是你要的月份缩写,赋值也会出错。
修正方案
方案1:修正循环逻辑
直接遍历月份缩写列表,检查哪个缩写在当前文件名里,找到后替换:
month_abv = ['Dec','Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov'] for inx in test.index: filename = test.loc[inx, 'month'] for abv in month_abv: if abv in filename: test.loc[inx, 'month'] = abv break # 找到匹配就停止,避免重复检查
方案2:更高效的向量化处理(推荐)
用Pandas的字符串提取功能,结合正则表达式直接提取月份缩写,比循环快得多:
month_abv = ['Dec','Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov'] # 把月份缩写拼成正则匹配模式 pattern = '|'.join(month_abv) # 提取匹配的内容 test['month'] = test['month'].str.extract(f'({pattern})', expand=False)
方案3:用apply函数处理
通过自定义函数对每个文件名做匹配,代码更简洁:
month_abv = ['Dec','Jan','Feb','Mar','Apr','May','Jun','Jul','Aug','Sep','Oct','Nov'] def extract_month(filename): for abv in month_abv: if abv in filename: return abv return filename # 没找到匹配时保留原内容 test['month'] = test['month'].apply(extract_month)
内容的提问来源于stack exchange,提问作者sahel
相关产品推荐
相关产品推荐

