You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将DataFrame中的二进制颜色标识列转换为对应的分类结果列

实现方案

我们可以通过Pandas快速实现该需求,无需硬编码颜色列名,适配任意数量的颜色列:

方案1:apply实现(易读易自定义规则)

该方法逻辑清晰,便于调整拼接规则,适合中小规模数据量:

import pandas as pd

# ------------- 此处替换为你自己的DataFrame即可 -------------
# 示例构造数据,实际使用时可以删掉这部分
df = pd.DataFrame({
    'black': [0,0,1],
    'orange': [1,0,0],
    'yellow': [0,0,0],
    'green': [1,1,0]
}, index=[1,2,3])

# 1. 定义所有颜色列,如果你整个df都是颜色列直接取df.columns即可
color_cols = df.columns.tolist()

# 2. 定义每行的拼接逻辑
def merge_color_names(row):
    # 筛选出值为1的对应颜色列名
    matched = [col for col, val in row.items() if val == 1]
    if len(matched) == 0:
        return "" # 无匹配颜色时返回空,可按需修改为其他默认值
    elif len(matched) == 1:
        return matched[0]
    else:
        # 2个及以上颜色时,最后一项前加and,超过2个时前面用逗号分隔
        return ", ".join(matched[:-1]) + " and " + matched[-1]

# 3. 生成新列
df["colours"] = df[color_cols].apply(merge_color_names, axis=1)

输出结果和要求完全一致:

black  orange  yellow  green          colours
1      0       1       0      1  orange and green
2      0       0       0      1            green
3      1       0       0      0            black

方案2:向量化运算(高效适配大数据量)

如果你的数据量/颜色列数量非常大,推荐使用该方案,性能比apply高几倍到几十倍:

# 筛选颜色列的步骤和方案1一致
color_cols = df.columns.tolist()

# 第一步:用矩阵乘法直接拼接所有匹配的颜色名
temp_series = df[color_cols].dot(pd.Index(color_cols) + ", ").str.rstrip(", ")

# 第二步:替换最后一个逗号为and
df["colours"] = temp_series.str.rsplit(", ", n=1).apply(lambda x: " and ".join(x) if len(x) > 1 else x[0])

内容的提问来源于stack exchange,提问作者FafamKurac123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 23:06:04