You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用自定义函数过滤Dask DataFrame?保留含男女法官的分组行

问题

我有一个包含以下列的DataFrame:

act     section     female_judge    time_taken
0   17.0    121384.0    0 nonfemale        321.0
1   423.0   273193.0    0 nonfemale        0.0
2   423.0   332490.0    0 nonfemale        0.0
3   481.0   1440.0      0 nonfemale        0.0
4   481.0   5145.0      0 nonfemale        0.0

我已按act、section和female_judge列对记录进行分组。现在我希望保留满足以下条件的行:对于给定的act和section,同时存在法官为女性和男性的记录(我想要对比相同act和section下男女法官的耗时情况)。
请问该如何过滤这些行?

解决方案

可以通过以下步骤实现:

  • 提取性别标识
    从数据格式来看,female_judge列包含性别描述,先将其转换为二元标识以便判断:

    df['is_female'] = df['female_judge'].str.contains('female').astype(int)
    
  • 筛选符合条件的组
    按act和section分组后,统计每组内的性别种类数,只保留同时包含男女的组:
    方法一:先获取有效组合再过滤

    # 计算每个(act, section)组的性别种类数,筛选出种类数为2的组
    valid_groups = df.groupby(['act', 'section'])['is_female'].nunique() == 2
    # 提取符合条件的(act, section)对
    valid_pairs = valid_groups[valid_groups].index
    # 过滤原DataFrame
    filtered_df = df[df.set_index(['act', 'section']).index.isin(valid_pairs)]
    

    方法二:用transform直接标记并过滤(更简洁)

    # 给每行标记其所在组是否同时有男女法官
    df['has_both_genders'] = df.groupby(['act', 'section'])['is_female'].transform(lambda x: x.nunique() == 2)
    # 筛选出符合条件的行
    filtered_df = df[df['has_both_genders']]
    
  • (可选)清理临时列
    如果不需要临时生成的is_female和has_both_genders列,可删除:

    filtered_df = filtered_df.drop(['is_female', 'has_both_genders'], axis=1)
    

处理后得到的filtered_df就是所有在相同act和section下同时存在男女法官记录的行,可直接用于耗时对比分析。

内容的提问来源于stack exchange,提问作者HaranCode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 10:45:24