You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将二进制列转换为基于列标题的字符串值列?

原始数据集
DeptCell cultureBioinfoImmunologyTrigonometryAlgebraMicrobioOptics
Biotech1110000
Biotech1010000
Math0001100
Biotech0000010
Physics0000001
目标转换结果
Dept课程1课程2课程3
BiotechCell cultureBioinfoImmunology
BiotechCell cultureImmunology
MathTrigonometryAlgebra
BiotechMicrobio
PhysicsOptics
实现方法

方法一:使用Pandas(推荐,适合大型数据集)

Pandas是Python处理表格数据的常用工具,代码简洁高效:

import pandas as pd

# 加载原始数据(实际场景可替换为pd.read_csv/pd.read_excel读取外部文件)
data = {
    "Dept": ["Biotech", "Biotech", "Math", "Biotech", "Physics"],
    "Cell culture": [1, 1, 0, 0, 0],
    "Bioinfo": [1, 0, 0, 0, 0],
    "Immunology": [1, 1, 0, 0, 0],
    "Trigonometry": [0, 0, 1, 0, 0],
    "Algebra": [0, 0, 1, 0, 0],
    "Microbio": [0, 0, 0, 1, 0],
    "Optics": [0, 0, 0, 0, 1]
}
df = pd.DataFrame(data)

# 提取每行值为1的课程名称
def extract_courses(row):
    return [col for col in df.columns[1:] if row[col] == 1]

df["Courses"] = df.apply(extract_courses, axis=1)

# 扩展课程列表为多列,空值用空字符串填充
max_course_count = df["Courses"].str.len().max()
for i in range(max_course_count):
    df[f"课程{i+1}"] = df["Courses"].str.get(i).fillna("")

# 生成最终结果
result = df[["Dept"] + [f"课程{i+1}" for i in range(max_course_count)]]

# 输出Markdown格式表格
print(result.to_markdown(index=False, colalign=["left"]*len(result.columns)))

方法二:纯Python实现(无需第三方库)

如果不想依赖外部库,可用原生Python处理列表:

# 原始数据
raw_data = [
    ["Dept", "Cell culture", "Bioinfo", "Immunology", "Trigonometry", "Algebra", "Microbio", "Optics"],
    ["Biotech", 1, 1, 1, 0, 0, 0, 0],
    ["Biotech", 1, 0, 1, 0, 0, 0, 0],
    ["Math", 0, 0, 0, 1, 1, 0, 0],
    ["Biotech", 0, 0, 0, 0, 0, 1, 0],
    ["Physics", 0, 0, 0, 0, 0, 0, 1]
]

headers = raw_data[0]
data_rows = raw_data[1:]

processed_rows = []
max_col_num = 1  # 至少包含Dept列

# 处理每一行,收集课程名称
for row in data_rows:
    dept = row[0]
    courses = []
    for col_idx, value in enumerate(row[1:]):
        if value == 1:
            courses.append(headers[col_idx + 1])
    processed_row = [dept] + courses
    if len(processed_row) > max_col_num:
        max_col_num = len(processed_row)
    processed_rows.append(processed_row)

# 填充空字符串,保证每行长度一致
for idx in range(len(processed_rows)):
    while len(processed_rows[idx]) < max_col_num:
        processed_rows[idx].append("")

# 生成Markdown表格
result_headers = ["Dept"] + [f"课程{i+1}" for i in range(max_col_num - 1)]
separator = ["---"] * max_col_num
markdown_lines = [result_headers, separator] + processed_rows

# 打印输出
for line in markdown_lines:
    print("| " + " | ".join(line) + " |")

内容的提问来源于stack exchange,提问作者Mimikyu o_0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 17:46:27