如何用正则表达式移除'-'及后续两位字符以提取采样日期?
提取采样日并转换为Pandas Datetime序列
下面提供两种代码实现方案,帮你快速处理数据并转换为datetime序列:
方案一:字符串分割处理
通过字符串分割提取所需的日、月、年部分,补全年份后转换格式:
import pandas as pd # 原始数据列表 raw_dates = [ "12-15.11.12", "19-22.11.12", "26-29.11.12", "03-06.12.12", "10-13.12.12", "17-20.12.12", "19-23.12.12", "27-30.12.12", "02-05.01.13" ] # 处理每个日期字符串 cleaned_dates = [] for s in raw_dates: # 拆分出采样日、后续的月年部分 sample_day, rest = s.split("-") month, year_short = rest.split(".")[1:] # 拼接成目标格式,补全20前缀的年份 cleaned = f"{sample_day}.{month}.20{year_short}" cleaned_dates.append(cleaned) # 转换为pandas datetime序列 datetime_series = pd.to_datetime(cleaned_dates, format="%d.%m.%Y") print(datetime_series)
方案二:正则表达式匹配
用正则直接提取关键部分,一步完成字符串替换:
import pandas as pd import re raw_dates = [ "12-15.11.12", "19-22.11.12", "26-29.11.12", "03-06.12.12", "10-13.12.12", "17-20.12.12", "19-23.12.12", "27-30.12.12", "02-05.01.13" ] # 正则匹配采样日、月、年,替换为目标格式 pattern = r"(\d{2})-\d{2}\.(\d{2})\.(\d{2})" cleaned_dates = [re.sub(pattern, r"\1.\2.20\3", s) for s in raw_dates] # 转换为datetime序列 datetime_series = pd.to_datetime(cleaned_dates, format="%d.%m.%Y") print(datetime_series)
输出结果示例
0 2012-11-12 1 2012-11-19 2 2012-11-26 3 2012-12-03 4 2012-12-10 5 2012-12-17 6 2012-12-19 7 2012-12-27 8 2013-01-02 dtype: datetime64[ns]
内容的提问来源于stack exchange,提问作者crtnnn
相关产品推荐
相关产品推荐

