You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编写自定义pandas aggfunc聚合后字段dtype全为object如何解决

问题背景

需要为geopandas.GeoDataFrame.dissolve()编写自定义聚合函数:合并多个多边形要素时,保留满足指定筛选条件、面积最大的多边形对应的全部属性信息。
当前自定义聚合逻辑可正常运行,但执行完成后GeoDataFrame的所有属性字段dtype均变为object类型,该问题在常规pandas groupby()聚合操作中同样可复现。
注意:无法使用convert_dtypes()方法修复该问题,因为geopandas.GeoDataFrame.dissolve()场景下convert_dtypes()会破坏几何列信息。

复现代码

import pandas as pd

df = pd.DataFrame({
    'group': ['A', 'A', 'B', 'B'],
    'ints': [1, 2, 3, 4],
    'floats': [1.0, 2.0, 2.2, 3.2],
    'strings': ['foo', 'bar', 'baz', 'qux'],
    'bools': [True, True, True, False],
    'test': ['drop this', 'keep this', 'keep this', 'drop this'],
    })


def custom_sort(df):
    """带自定义排序规则的聚合函数"""
    df = df.sort_values(by=['bools', 'floats'], ascending=False)
    return df.iloc[0]


print(df)
print(df.dtypes)
print()
grouped = df.groupby(by='group').agg(custom_sort)
print(grouped)
print(grouped.dtypes)  # 问题:所有字段dtype都变为object
print()
print(grouped.convert_dtypes().dtypes)  # 该方案不适用geopandas场景

# 无法使用convert_dtypes(),geopandas.GeoDataFrame.dissolve()场景下该方法会破坏几何列信息

执行输出

group  ints  floats strings  bools       test
0     A     1     1.0     foo   True  drop this
1     A     2     2.0     bar   True  keep this
2     B     3     2.2     baz   True  keep this
3     B     4     3.2     qux  False  drop this
group       object
ints         int64
floats     float64
strings     object
bools         bool
test        object
dtype: object

      ints floats strings bools       test
group                                     
A        2    2.0     bar  True  keep this
B        3    2.2     baz  True  keep this
ints       object
floats     object
strings    object
bools      object
test       object
dtype: object

ints         Int64
floats     Float64
strings     string
bools      boolean
test        string
dtype: object
原因说明

自定义聚合函数custom_sort()返回的是包含多类型值的单行Series,pandas处理这类跨类型Series时会自动将整体dtype降级为object,聚合拼接后所有字段都会被统一转为object类型。

解决方法

不要使用返回整行Series的聚合函数传入agg(),直接通过排序+分组取首行索引的方式筛选目标行,全程不会触发跨类型转换,可完整保留各字段原有dtype,且兼容geopandas几何列:

# 先按分组键、自定义排序规则排序
sorted_df = df.sort_values(
    by=['group', 'bools', 'floats'],
    ascending=[True, False, False]
)
# 取每个分组的第一行,重置索引即可
grouped = sorted_df.groupby('group').head(1).set_index('group')

# 验证dtype
print(grouped.dtypes)

执行后输出的dtype与原表完全一致:

ints         int64
floats     float64
strings     object
bools         bool
test        object
dtype: object

如果是geopandas dissolve场景,可先单独对几何列做分组合并,再用上述方法筛出对应属性后拼接,不会破坏几何列信息,也不会出现dtype异常。

内容的提问来源于stack exchange,提问作者Azrael_DD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 10:06:37