Pandas分组最小时间戳与分组时间戳比较报错排查
Pandas分组比较触发RecursionError的解决方案
Hey there, let's break down what's causing that RecursionError and fix your code to get the desired new_col!
Your Scenario Recap
You're working with a subset of a DataFrame where id3 = 1041, and you want to:
- Get the minimum value of
max_snsr_tsgrouped byid3 - Compare that value to the grouped
max_ts_fs(grouped byid3andk) - Add a boolean column
new_colbased on the comparison
What's Wrong with Your Original Code?
Let's look at the problematic line:
joined_h_raw_fs['new_col'] = np.where(joined_h_raw_fs.groupby(['id3'])['max_snsr_ts'].min().min() > joined_h_raw_fs.groupby(['id3', 'k'])['max_ts_fs'] , True, False)
There are two key issues here:
- Redundant & Incorrect Aggregation:
groupby(['id3'])['max_snsr_ts'].min()already returns a Series with one value perid3(for your subset, it's just the min forid3=1041). Adding an extra.min()collapses this to a single scalar value, which breaks alignment with your DataFrame rows. - Mismatched Indexes & Unmapped Grouped Values: When you compare a scalar/grouped Series directly to another grouped Series, Pandas tries to align indexes across mismatched structures. This, combined with potential unparsed timestamp strings (instead of proper datetime objects), triggers the
RecursionErrorduring comparison.
Step-by-Step Fix
First, make sure your timestamp columns are parsed as proper datetime objects (this avoids string-comparison quirks that can cause errors):
import pandas as pd import numpy as np # Convert string timestamps to datetime objects joined_h_raw_fs['max_snsr_ts'] = pd.to_datetime(joined_h_raw_fs['max_snsr_ts']) joined_h_raw_fs['max_ts_fs'] = pd.to_datetime(joined_h_raw_fs['max_ts_fs'])
Next, use transform() to map grouped aggregation results back to every row of your original DataFrame (this keeps indexes aligned correctly):
# Get the min max_snsr_ts for each id3, mapped to every row id3_min_snsr = joined_h_raw_fs.groupby('id3')['max_snsr_ts'].transform('min') # Get the max max_ts_fs for each (id3, k) group, mapped to every row id3_k_max_ts = joined_h_raw_fs.groupby(['id3', 'k'])['max_ts_fs'].transform('max') # Create the boolean column joined_h_raw_fs['new_col'] = id3_min_snsr > id3_k_max_ts
What This Does
transform()ensures that each row gets the aggregated value corresponding to its group, so you're comparing values row-by-row instead of trying to align mismatched grouped Series.- Parsing timestamps to datetime objects ensures accurate time-based comparisons, not string dictionary-order comparisons.
Result for Your Sample Data
For your subset where id3=1041:
- The minimum
max_snsr_tsis2020-10-19 23:59:00 - For
k=48, the groupedmax_ts_fsis2020-10-22 23:30:00 - For
k=96, the groupedmax_ts_fsis2020-10-23 23:30:00
All comparisons will returnFalse, sonew_colwill beFalsefor every row, which aligns with the logic you want.
内容的提问来源于stack exchange,提问作者tfkLSTM
相关产品推荐
相关产品推荐

