Python Pandas如何筛选1988-2018年数据并存入selection变量
实现方法
你已经将日期列转为DatetimeIndex并设为行索引,直接使用pandas针对时间索引的原生切片语法即可完成筛选,不需要逐行判断日期。
核心筛选代码仅需1行:
selection = df.loc['1988-01':'2018-12']
说明:pandas对已排序的DatetimeIndex支持直接用
'年-月'格式的字符串做闭区间切片,上述代码会自动覆盖1988年1月1日到2018年12月31日的所有记录,无需手动指定当月最后一天的精确日期。
注意事项
- 如果读取文件后日期索引不是升序排列,先执行排序再切片,避免筛选结果缺失:
df = df.sort_index() selection = df.loc['1988-01':'2018-12']
- 筛选完成后可以通过以下代码核对时间范围是否符合要求:
# 打印筛选后数据的最早、最晚日期 print("筛选时段起始日期:", selection.index.min()) print("筛选时段结束日期:", selection.index.max()) # 打印筛选后数据行数 print("筛选后数据总行数:", len(selection))
调整后完整参考代码
import pandas as pd import pandas_datareader as pdr import matplotlib.pyplot as plt from datetime import date df = pd.read_csv('helsinki-vantaa.csv.csv', parse_dates=['DATE'], index_col=['DATE']) # 确保时间索引升序排列 df = df.sort_index() # 基础信息查看 df.head() rows_count = df.shape[0] print("原始数据总行数:", rows_count) # 筛选1988年1月-2018年12月数据 selection = df.loc['1988-01':'2018-12'] # 验证筛选结果 print("筛选后最早日期:", selection.index.min()) print("筛选后最晚日期:", selection.index.max()) print("筛选后数据总行数:", len(selection))
内容的提问来源于stack exchange,提问作者Jada Parslow
相关产品推荐
相关产品推荐

