You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas Groupby Apply计算月度消费比率时触发索引不兼容TypeError

月度消费比率计算报错原因分析

执行代码water_df["consumption_ratio"] = water_df.groupby(['Datetime', 'houseid-meterid']).apply(consumption_ratio)时触发TypeError,提示“incompatible index of inserted column with frame index”,以下是问题成因分析:

消费比率函数代码

def consumption_ratio(row): 
    c_consumption = row["consumption"].iloc[0]
    month = row["month"].iloc[0]
    year = row["year"].iloc[0]
    house = row["houseid-meterid"].iloc[0]

    if month == 2 and year == 2019: 
        return 0
    else: 
        if month == 1:
            prevyear = year - 1
            prevmonth = 12
            prev_record = water_df.query("`houseid-meterid` == @house and year == @prevyear and month == @prevmonth")
            try:
                ratio = c_consumption / prev_record["consumption"]
            except ZeroDivisionError:
                ratio = 0
            return ratio
        else: 
            prevmonth  = month - 1
            prev_record = water_df.query("`houseid-meterid` == @house and year == @year and month == @prevmonth")
            try:
                ratio = c_consumption/ prev_record["consumption"]
            except ZeroDivisionError:
                ratio = 0
            return ratio

报错栈

ValueError                                Traceback (most recent call last)
File D:\ML Projects\Bityarn-UtilitiesAnalysis\venv\lib\site-packages\pandas\core\frame.py:12017, in _reindex_for_setitem(value, index)
  12016 try:
> 12017     reindexed_value = value.reindex(index)._values
  12018 except ValueError as err:
  12019     # raised in MultiIndex.from_tuples, see test_insert_error_msmgs

...(省略中间栈信息)

TypeError: incompatible index of inserted column with frame index

数据集结构

year    month   houseid-meterid     Datetime    cleaned_quantity
2019    2      m5             2019-02-01              7.0
2019    3      m5             2019-03-01              23.0
2019    4      m5             2019-04-01              14.0
2019    4      m5             2019-05-01              22.0

预期输出

year    month   houseid-meterid     Datetime     consumption  consumption-ratio
2019    2      m5             2019-02-01              7.0           0
2019    3      m5             2019-03-01              23.0          3.285
2019    4      m5             2019-04-01              14.0          0.608
2019    4      m5             2019-05-01              22.0          1.571

问题成因

  1. 分组索引不匹配
    用['Datetime', 'houseid-meterid']作为分组键时,每个分组仅包含单条记录。groupby.apply()返回的结果会保留分组的MultiIndex(由Datetime和houseid-meterid组成),但原DataFrame使用默认整数索引,两者结构不一致,插入新列时触发索引冲突。

  2. 函数返回值类型异常
    若prev_record包含多条记录,prev_record["consumption"]是Series,此时c_consumption / prev_record["consumption"]返回带索引的Series而非单个标量,导致apply()返回结果的维度与原DataFrame不匹配,加剧索引冲突。

  3. 全局变量引用风险
    函数内部直接引用全局的water_df,而非基于分组上下文处理数据,破坏封装性的同时,可能导致查询结果索引与分组索引冲突,间接引发索引不兼容问题。

内容的提问来源于stack exchange,提问作者Shangazi Mkubwa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 17:51:27