You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas如何对分类特征相同的行执行非重复凸组合运算

实现代码

首先导入依赖库:

import pandas as pd
from itertools import combinations

步骤1:定义参数与数据准备

# 构造示例数据,可替换为自己的DataFrame
df = pd.DataFrame({
    "Is taco?": pd.Categorical(["yes", "yes", "yes"]),
    "Count": [2, 5, 1]
})

# 凸组合参数L,可根据需求调整,取值范围[0,1]
L = 0.5

# 自动识别分类列与数值列
cat_cols = df.select_dtypes(include="category").columns.tolist()
num_cols = df.select_dtypes(include="number").columns.tolist()

步骤2:分组计算凸组合

result = []
# 按所有分类列分组,仅同分类组内的行做配对
for _, group in df.groupby(cat_cols, observed=True):
    # 生成组内不重复的两两行组合(i<j,天然避免重复配对)
    for (_, row1), (_, row2) in combinations(group.iterrows(), 2):
        # 分类特征直接继承原值,数值特征做凸组合计算
        new_row = pd.concat([
            row1[cat_cols],
            L * row1[num_cols] + (1 - L) * row2[num_cols]
        ])
        result.append(new_row)

# 转换为DataFrame并重置索引
final_df = pd.DataFrame(result).reset_index(drop=True)

输出验证

运行上述代码后打印final_df即可得到预期结果:

Is taco?  Count
0      yes    3.5
1      yes    1.5
2      yes    3.0

特性说明

  • 自动支持多分类列:只要是dtype为category的列都会作为分组依据,只有所有分类列取值完全一致的行才会配对
  • 自动支持多数值列:所有数值类型的列都会自动做凸组合计算,无需单独修改逻辑
  • 无重复配对:使用combinations生成无序两两组合,保证任意两行仅计算一次
  • 可灵活调整凸组合系数:修改L的取值即可适配不同的凸组合需求

内容的提问来源于stack exchange,提问作者user12314098

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 14:24:04