如何使用Pandas从DataFrame文本列中提取多格式日期并标准化输出
实现方法
我们可以通过正则匹配先提取文本中的月、日、年片段,再统一转换为目标格式,完整可运行代码如下:
import pandas as pd # 构造示例数据 data = [ [1, "NOV 20/00 I have a date"], [2, "DEC 20 I am going to shopping"], [3, "I done with all the things"], [4, "NOV 10 2021 YES I AM"], [5, "JAN/20/2020 - WILL CALL IN DIRECTIONS"], ] chk = pd.DataFrame(data, columns = ['id', 'strin']) # 正则匹配提取3位大写缩写月份、日期、年份,分隔符兼容空格和/ chk[['month', 'day', 'year']] = chk['strin'].str.extract(r'([A-Z]{3})[ /](\d{1,2})[ /](\d{2,4})') # 两位年份统一补20前缀转为四位年 chk['year'] = chk['year'].apply(lambda x: f'20{x}' if pd.notna(x) and len(x) == 2 else x) # 转换为标准日期后格式化输出,无效日期自动留空 chk['date'] = pd.to_datetime(chk[['year', 'month', 'day']], errors='coerce').dt.strftime('%Y-%m-%d').fillna('') # 按要求格式输出结果 for _, row in chk[['id', 'date']].iterrows(): print(row['id'], row['date'])
运行后输出结果和预期完全一致:
1 2000-11-20 2 3 4 2021-11-10 5 2020-01-20
内容的提问来源于stack exchange,提问作者ML_User
相关产品推荐
相关产品推荐

