You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas .last()返回数据超出预期,是用法问题还是库问题?

Pandas .last()方法筛选日期数据不符合预期的问题分析

问题描述

使用Pandas的.last()方法筛选日期数据时,指定7D、14D等时间范围返回的数据超出预期:例如CAT股票近一周无交易记录,但代码却返回了一条两周前的记录,使用1W参数也得到相同结果。

获取数据代码

# Gets the openinsider information for insider trading of stocks
def get_openinsider_stock_activity(stocks):
    # Format search string from stocks
    # URL gets 500 entries only
    stocks = stocks.upper().replace(' ', '+')  # format the tickers to be lis
    print('Search String: ', stocks)
    url = f'http://openinsider.com/screener?s={stocks}&o=&pl=&ph=&ll=&lh=&fd=1461&fdr=&td=0&tdr=&fdlyl=&fdlyh=&daysago=&xp=1&xs=1&vl=&vh=&ocl=&och=&sic1=-1&sicl=100&sich=9999&grp=0&nfl=&nfh=&nil=&nih=&nol=&noh=&v2l=&v2h=&oc2l=&oc2h=&sortcol=0&cnt=300&page=1'
    response = requests.get(url)
    soup = BeautifulSoup(response.text, 'lxml')
    try:
        rows = soup.find('table', {'class': 'tinytable'}).find('tbody').findAll('tr')
    except Exception as e:
        print("Error! Skipping")
        print(f"This URL was not successful: {url}")
        print(e)
        return

    # Get headers first
    header_row = soup.find('table', {'class': 'tinytable'}).find('thead').findAll('tr')
    headers = []
    for row in header_row:
        for col in row:
            for header in col:
                text = col.text.strip()
                if text == '':
                    continue
                headers.append(unicodedata.normalize('NFKD', text))

    body_rows = soup.find('table', {'class': 'tinytable'}).find('tbody').findAll('tr')
    # Iterate over rows of table
    insider_data = []
    for row in body_rows:
        cols = row.findAll('td')
        if not cols:
            continue
        for cell in cols:
            cell_value = cell.find('a').text.strip() if cell.find('a') else cell.text.strip()
        body = {key: cols[index].find('a').text.strip() if cols[index].find('a') else cols[index].text.strip()
                for index, key in enumerate(headers)}
        insider_data.append(body)
    insider_data = pd.DataFrame(insider_data)
    return insider_data

筛选数据代码(原问题代码)

data = []
for stock in ['AAPL', 'BA', 'GME', 'XOM', 'CAT']:
    data.append(get_openinsider_stock_activity(stock))
data = pd.concat(data)

for timeframe in ['7D', '14D', '1M']:

    pd.set_option('display.max_columns', 500)
    pd.set_option('display.width', 1000)
    pd.set_option('display.max_rows', None)

    count_data = data.copy()
    count_data['Trade Date'] = pd.to_datetime(count_data['Trade Date'])
    # 问题出在这里
    count_data = count_data.sort_values(by='Trade Date', ascending=True).set_index('Trade Date').last(timeframe)
    print(count_data)

问题原因

这不是Pandas的bug,是对.last()方法的用法理解错误:

  • .last()方法的作用是从整个数据集的最后一个时间点往前倒推指定时长,筛选这个时间范围内的记录,而不是从当前系统时间往前倒推筛选最近N天的数据。
  • 以CAT股票为例,数据集里最后一条记录是两周前的,使用.last('7D')时,会以这条记录的日期为终点,往前取7天内的记录,这条两周前的记录自然会被包含在内,但实际需要的是相对于今天的7天内的记录。

修正方案

改用基于当前时间的布尔索引来筛选最近N天的数据,具体代码替换如下:

data = []
for stock in ['AAPL', 'BA', 'GME', 'XOM', 'CAT']:
    data.append(get_openinsider_stock_activity(stock))
data = pd.concat(data)

for timeframe in ['7D', '14D', '1M']:

    pd.set_option('display.max_columns', 500)
    pd.set_option('display.width', 1000)
    pd.set_option('display.max_rows', None)

    count_data = data.copy()
    count_data['Trade Date'] = pd.to_datetime(count_data['Trade Date'])
    # 计算时间阈值:当前时间减去指定时长
    threshold = pd.Timestamp.now() - pd.Timedelta(timeframe)
    # 筛选Trade Date >= 阈值的记录
    count_data = count_data[count_data['Trade Date'] >= threshold]
    # 可选:按日期排序
    count_data = count_data.sort_values(by='Trade Date', ascending=True)
    print(count_data)

补充说明

  • 如果需要忽略时间部分(只比较日期),可以把pd.Timestamp.now()换成pd.Timestamp.today().normalize()。
  • 若数据中存在未来日期(爬取错误导致),可额外添加count_data['Trade Date'] <= pd.Timestamp.now()的条件过滤。

内容的提问来源于stack exchange,提问作者David Frick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 14:10:35