Pandas中FutureWarning警告及结果展示异常求助
Pandas日期提取排序异常及警告问题解决
问题背景
作为Pandas和正则表达式新手,需求为从文件提取指定格式日期,按时间排序后添加原索引列,通过主代码调用函数实现。目前函数逻辑看似正确,但运行时出现FutureWarning警告,且主函数仅展示index列。
问题原因
1. 主函数仅展示index列
- 主代码中使用
print(result.T.head(1).T),该操作先将DataFrame转置,取第一行(对应原DataFrame的index列)后再转置,导致仅输出index列。 - 函数定义为
date_sort,但主代码调用date_sorter(),存在函数名拼写错误(从输出推断应为笔误)。
2. FutureWarning警告
- 使用
df['year'].str.extract(r'(\d?\d\d\d)')时未指定expand参数,当前Pandas版本中expand=None等价于expand=False,但未来版本会默认改为expand=True,因此触发警告。
解决方法与优化代码
针对主函数输出问题
- 直接打印
result或result.head()即可查看完整的原索引和年份列,移除多余的转置操作。 - 修正函数名,确保定义与调用一致。
针对FutureWarning警告
- 在
str.extract中显式指定expand=False(提取单个组时返回Series,更简洁)或expand=True(返回DataFrame)。
其他优化
- 将日期列表作为参数传入函数,避免依赖全局变量,提升代码规范性。
- 修正年份提取正则为
(\d{4}),精准匹配4位年份。 - 将年份转换为整数类型后排序,避免字符串排序的潜在问题。
- 删除无实际作用的
df2['year'].str.len()代码行。
修正后完整代码
import re import pandas as pd def date_sort(date_list): df = pd.DataFrame(date_list, columns=['full_date']) # 提取4位年份,显式指定expand=False返回Series df['year'] = df['full_date'].str.extract(r'(\d{4})', expand=False) # 转换年份为整数类型,确保排序逻辑正确 df['year'] = df['year'].astype(int) # 按年份排序,重置索引并保留原索引为新列 df_sorted = df.sort_values('year').reset_index(names='original_index') # 返回包含原索引、完整日期、年份的结果 return df_sorted[['original_index', 'full_date', 'year']] # 主代码执行入口 if __name__ == "__main__": with open('text_with_lot_of_text_and_dates.txt', 'r') as file: tot6 = file.read() # 提取指定格式的日期字符串 date_reader = re.findall(r'(?:\d{1,2} )?(?:Jan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Nov|Dec)[a-z]* (?:\d{1,2}, )?\d{4}', tot6) result = date_sort(date_reader) # 打印前5行结果 print(result.head())
预期输出
运行后将输出包含原索引、完整日期、年份的DataFrame示例:
original_index full_date year 0 4 06 May 1972 1972 1 21 13 Jan 1972 1972 2 47 [对应日期] 1972 3 72 [对应日期] 1972 4 162 [对应日期] 1973
内容的提问来源于stack exchange,提问作者Juan Andres Mercado
相关产品推荐
相关产品推荐

