使用pd.cut分组DataFrame时出现Grouper与轴长度不一致的ValueError求助
Ah, I’ve run into this exact gotcha before! Let’s figure out what’s going wrong here.
The Root Cause
When you use pd.cut() with the retbins=True parameter, it doesn’t just return your binned labels—it returns a tuple containing two items:
- The binned label Series (which matches the length of your original
df['time']column) - The actual bin boundaries array (which has length 8 in your case, since you defined 7 bins)
When you pass this entire tuple directly to groupby(), Pandas gets confused because the second element of the tuple (the bins array) has a different length than your DataFrame’s rows. That’s exactly why you’re seeing the ValueError: Grouper and axis must be same length error.
The Fixes
There are two simple ways to resolve this, depending on whether you need to keep the returned bin boundaries or not:
1. If you need to retain the bin boundaries
Split the tuple returned by pd.cut() into separate variables, so you only pass the binned labels to groupby():
# Split the returned tuple into binned data and bins frameddata, returned_bins = pd.cut(df['time'], bins, retbins=True, labels=timeframe) # Now group using only the frameddata Series groups = df.groupby(frameddata)
2. If you don’t need the bin boundaries
Just remove the retbins=True parameter entirely. pd.cut() will then only return the binned label Series, which matches your DataFrame’s length perfectly:
# No retbins=True, so frameddata is just the binned labels frameddata = pd.cut(df['time'], bins, labels=timeframe) # Grouping works as expected groups = df.groupby(frameddata)
Quick Bonus Tip
If your df['time'] has values outside the [3, 24] range, pd.cut() will default to setting those values to NaN. If you want to include these NaN values as a separate group in your results, add the dropna=False parameter to groupby():
groups = df.groupby(frameddata, dropna=False)
内容的提问来源于stack exchange,提问作者user3194861

