You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame中高效将C、D合并为'NOT A and B'分组?

优化Pandas DataFrame类别合并的实现方式

针对你要把C、D类别合并为"NOT A and B"并求和的需求,这里有几种比逐个替换更高效的实现方式:

方法1:用numpy.where映射后分组

不用提前修改原数据,直接在分组时完成类别映射,代码简洁高效:

import pandas as pd
import numpy as np

# 构造原DataFrame
df = pd.DataFrame({
    'Category': ['A', 'B', 'C', 'D'],
    'count': [327, 20, 30, 302]
})

# 映射类别并分组求和
result = df.groupby(
    np.where(df['Category'].isin(['A', 'B']), df['Category'], 'NOT A and B'),
    as_index=False
)['count'].sum()

# 重置列名
result.columns = ['Category', 'count']

方法2:拆分+拼接,逻辑直观

先提取A、B的原始行,单独计算其他类别的总和,再拼接结果,适合需要单独处理不同组的场景:

# 提取A、B的行
ab_rows = df[df['Category'].isin(['A', 'B'])]
# 计算C、D的总和
other_total = pd.DataFrame({
    'Category': ['NOT A and B'],
    'count': [df[~df['Category'].isin(['A', 'B'])]['count'].sum()]
})
# 拼接成最终结果
result = pd.concat([ab_rows, other_total], ignore_index=True)

方法3:简化版replace+分组

如果习惯用replace,可以直接用字典一次性映射,不用逐个替换C、D:

# 定义映射规则
category_map = {'C': 'NOT A and B', 'D': 'NOT A and B'}
# 替换后分组求和
result = df.replace({'Category': category_map}).groupby('Category', as_index=False)['count'].sum()

这些方法的优势:

  • 方法1和3无需拆分DataFrame,代码更紧凑,处理大数据量时效率更高;
  • 方法2逻辑清晰,容易理解和调整,适合需要对不同组做额外处理的情况。

内容的提问来源于stack exchange,提问作者user16971617

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:35:15