如何用正则表达式从多DataFrame中提取HE开头的时段数值
解决正则表达式提取DataFrame中所有HE开头时段数值的问题
问题原因
当前代码只提取到第一个HE开头的数值,是因为默认的正则匹配方法(如re.search或str.extract)仅返回首个匹配项。要获取所有符合条件的结果,需要使用支持全局匹配的方法。
解决方案
以下两种方法均可提取所有HE开头的时段数值,并整理为你需要的元组格式:
方法一:结合re.findall与apply批量处理
import pandas as pd import re Example_dict = {'S_1':pd.DataFrame({'Product1 Hours':'Lowest priced 5 concecutive hours in HE09 thru HE16', 'Product2 Hours': 'Highest priced 4 consecutive hours in HE17 through HE22'},index=[1,2]).T.reset_index(), 'S_2':pd.DataFrame({'Product1 Hours':'Lowest priced 5 concecutive hours in HE09 thru HE16', 'Product2 Hours': 'Highest priced 4 consecutive hours in HE17 through HE23'},index=[1,2]).T.reset_index() } # 定义提取函数:捕获所有HE后的两位数字,按每两个一组转为元组 def get_he_periods(text): hour_matches = re.findall(r'HE(\d{2})', text) return [tuple(hour_matches[i:i+2]) for i in range(0, len(hour_matches), 2)] # 遍历字典处理每个DataFrame for contract, df in Example_dict.items(): df['Extracted Periods'] = df[1].apply(get_he_periods) print(f"=== {contract} 处理结果 ===") print(df[['index', 'Extracted Periods']])
方法二:使用Pandas原生str.extractall
for contract, df in Example_dict.items(): # 提取所有HE后的数字,按行分组转为列表 all_matches = df[1].str.extractall(r'HE(\d{2})')[0].groupby(level=0).agg(list) # 转换为时段元组对 df['Extracted Periods'] = all_matches.apply(lambda x: [tuple(x[i:i+2]) for i in range(0, len(x), 2)]) print(f"=== {contract} 处理结果 ===") print(df[['index', 'Extracted Periods']])
输出结果
- S_1的
Product1 Hours会得到[('09', '16')],Product2 Hours得到[('17', '22')] - S_2的
Product2 Hours会得到[('17', '23')],完全符合预期需求
核心说明
re.findall(r'HE(\d{2})', text):全局匹配所有HE开头的两位数字,返回匹配到的数字列表- 切片分组
[tuple(x[i:i+2])...]:将连续的两个数字组合为一个时段元组 str.extractall:Pandas专门用于提取所有匹配项的方法,适合DataFrame列的批量处理
内容的提问来源于stack exchange,提问作者ARE
相关产品推荐
相关产品推荐

