如何将二进制列转换为基于列标题的字符串值列?
原始数据集
| Dept | Cell culture | Bioinfo | Immunology | Trigonometry | Algebra | Microbio | Optics |
|---|---|---|---|---|---|---|---|
| Biotech | 1 | 1 | 1 | 0 | 0 | 0 | 0 |
| Biotech | 1 | 0 | 1 | 0 | 0 | 0 | 0 |
| Math | 0 | 0 | 0 | 1 | 1 | 0 | 0 |
| Biotech | 0 | 0 | 0 | 0 | 0 | 1 | 0 |
| Physics | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
目标转换结果
| Dept | 课程1 | 课程2 | 课程3 |
|---|---|---|---|
| Biotech | Cell culture | Bioinfo | Immunology |
| Biotech | Cell culture | Immunology | |
| Math | Trigonometry | Algebra | |
| Biotech | Microbio | ||
| Physics | Optics |
实现方法
方法一:使用Pandas(推荐,适合大型数据集)
Pandas是Python处理表格数据的常用工具,代码简洁高效:
import pandas as pd # 加载原始数据(实际场景可替换为pd.read_csv/pd.read_excel读取外部文件) data = { "Dept": ["Biotech", "Biotech", "Math", "Biotech", "Physics"], "Cell culture": [1, 1, 0, 0, 0], "Bioinfo": [1, 0, 0, 0, 0], "Immunology": [1, 1, 0, 0, 0], "Trigonometry": [0, 0, 1, 0, 0], "Algebra": [0, 0, 1, 0, 0], "Microbio": [0, 0, 0, 1, 0], "Optics": [0, 0, 0, 0, 1] } df = pd.DataFrame(data) # 提取每行值为1的课程名称 def extract_courses(row): return [col for col in df.columns[1:] if row[col] == 1] df["Courses"] = df.apply(extract_courses, axis=1) # 扩展课程列表为多列,空值用空字符串填充 max_course_count = df["Courses"].str.len().max() for i in range(max_course_count): df[f"课程{i+1}"] = df["Courses"].str.get(i).fillna("") # 生成最终结果 result = df[["Dept"] + [f"课程{i+1}" for i in range(max_course_count)]] # 输出Markdown格式表格 print(result.to_markdown(index=False, colalign=["left"]*len(result.columns)))
方法二:纯Python实现(无需第三方库)
如果不想依赖外部库,可用原生Python处理列表:
# 原始数据 raw_data = [ ["Dept", "Cell culture", "Bioinfo", "Immunology", "Trigonometry", "Algebra", "Microbio", "Optics"], ["Biotech", 1, 1, 1, 0, 0, 0, 0], ["Biotech", 1, 0, 1, 0, 0, 0, 0], ["Math", 0, 0, 0, 1, 1, 0, 0], ["Biotech", 0, 0, 0, 0, 0, 1, 0], ["Physics", 0, 0, 0, 0, 0, 0, 1] ] headers = raw_data[0] data_rows = raw_data[1:] processed_rows = [] max_col_num = 1 # 至少包含Dept列 # 处理每一行,收集课程名称 for row in data_rows: dept = row[0] courses = [] for col_idx, value in enumerate(row[1:]): if value == 1: courses.append(headers[col_idx + 1]) processed_row = [dept] + courses if len(processed_row) > max_col_num: max_col_num = len(processed_row) processed_rows.append(processed_row) # 填充空字符串,保证每行长度一致 for idx in range(len(processed_rows)): while len(processed_rows[idx]) < max_col_num: processed_rows[idx].append("") # 生成Markdown表格 result_headers = ["Dept"] + [f"课程{i+1}" for i in range(max_col_num - 1)] separator = ["---"] * max_col_num markdown_lines = [result_headers, separator] + processed_rows # 打印输出 for line in markdown_lines: print("| " + " | ".join(line) + " |")
内容的提问来源于stack exchange,提问作者Mimikyu o_0
相关产品推荐
相关产品推荐

