You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组最小时间戳与分组时间戳比较报错排查

Pandas分组比较触发RecursionError的解决方案

Hey there, let's break down what's causing that RecursionError and fix your code to get the desired new_col!

Your Scenario Recap

You're working with a subset of a DataFrame where id3 = 1041, and you want to:

  • Get the minimum value of max_snsr_ts grouped by id3
  • Compare that value to the grouped max_ts_fs (grouped by id3 and k)
  • Add a boolean column new_col based on the comparison

What's Wrong with Your Original Code?

Let's look at the problematic line:

joined_h_raw_fs['new_col'] = np.where(joined_h_raw_fs.groupby(['id3'])['max_snsr_ts'].min().min() > joined_h_raw_fs.groupby(['id3', 'k'])['max_ts_fs'] , True, False)

There are two key issues here:

  1. Redundant & Incorrect Aggregation: groupby(['id3'])['max_snsr_ts'].min() already returns a Series with one value per id3 (for your subset, it's just the min for id3=1041). Adding an extra .min() collapses this to a single scalar value, which breaks alignment with your DataFrame rows.
  2. Mismatched Indexes & Unmapped Grouped Values: When you compare a scalar/grouped Series directly to another grouped Series, Pandas tries to align indexes across mismatched structures. This, combined with potential unparsed timestamp strings (instead of proper datetime objects), triggers the RecursionError during comparison.

Step-by-Step Fix

First, make sure your timestamp columns are parsed as proper datetime objects (this avoids string-comparison quirks that can cause errors):

import pandas as pd
import numpy as np

# Convert string timestamps to datetime objects
joined_h_raw_fs['max_snsr_ts'] = pd.to_datetime(joined_h_raw_fs['max_snsr_ts'])
joined_h_raw_fs['max_ts_fs'] = pd.to_datetime(joined_h_raw_fs['max_ts_fs'])

Next, use transform() to map grouped aggregation results back to every row of your original DataFrame (this keeps indexes aligned correctly):

# Get the min max_snsr_ts for each id3, mapped to every row
id3_min_snsr = joined_h_raw_fs.groupby('id3')['max_snsr_ts'].transform('min')

# Get the max max_ts_fs for each (id3, k) group, mapped to every row
id3_k_max_ts = joined_h_raw_fs.groupby(['id3', 'k'])['max_ts_fs'].transform('max')

# Create the boolean column
joined_h_raw_fs['new_col'] = id3_min_snsr > id3_k_max_ts

What This Does

  • transform() ensures that each row gets the aggregated value corresponding to its group, so you're comparing values row-by-row instead of trying to align mismatched grouped Series.
  • Parsing timestamps to datetime objects ensures accurate time-based comparisons, not string dictionary-order comparisons.

Result for Your Sample Data

For your subset where id3=1041:

  • The minimum max_snsr_ts is 2020-10-19 23:59:00
  • For k=48, the grouped max_ts_fs is 2020-10-22 23:30:00
  • For k=96, the grouped max_ts_fs is 2020-10-23 23:30:00
    All comparisons will return False, so new_col will be False for every row, which aligns with the logic you want.

内容的提问来源于stack exchange,提问作者tfkLSTM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:17:57