You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何转换仅含0和1值的多列二进制编码数据集?

Hey there! Let's break down how to transform your binary indicator DataFrame (with columns V1 to V21 containing only 0s and 1s) into a more usable format. This type of data is typically called one-hot encoded (if each sample has exactly one 1) or a multi-label binary dataset (if samples can have multiple 1s). Here are the most common solutions based on your end goal:

1. Convert to a Single Categorical Column (Single-Label Scenario)

If each row has exactly one 1 (meaning each sample belongs to one category), you can map the binary columns to a single category label:

import pandas as pd

# Structure your sample data into a proper DataFrame
sample_data = [
    [1, 0, 0, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0, 1, 0],
    [2, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1],
    [3, 0, 0, 0, 1, 1, 0, 0, 0, 0, 1, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0],
    [4, 0, 0, 0, 1, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0, 0, 1, 0, 1, 0, 0]
]

df = pd.DataFrame(sample_data, columns=["Sample"] + [f"V{i}" for i in range(1, 22)])

# Get the column name where the value is 1 for each row
df["Category"] = df.filter(regex="V\d+").idxmax(axis=1)

# Optional: Extract just the numeric part of the category (e.g., "V5" → 5)
df["Category_ID"] = df["Category"].str.extract(r"(\d+)").astype(int)

This gives you a new column with the category label (like V4 or 5) corresponding to the single active feature in each row.

2. Convert to Multi-Label Lists (Multi-Label Scenario)

If rows can have multiple 1s (meaning each sample belongs to multiple categories), collect all active feature labels into a list:

# For each row, gather all V columns where the value is 1
df["Active_Labels"] = df.filter(regex="V\d+").apply(
    lambda row: row[row == 1].index.tolist(),
    axis=1
)

# Optional: Convert labels to numeric IDs instead of "Vx" strings
df["Active_Label_IDs"] = df.filter(regex="V\d+").apply(
    lambda row: [int(col[1:]) for col in row[row == 1].index],
    axis=1
)

This results in a column where each entry is a list of all categories the sample belongs to (e.g., ['V4', 'V8', 'V10'] or [4,8,10]).

3. Create a Compressed Binary String (For Feature Engineering)

If you just need to condense the 21 binary columns into a single feature, convert each row's 0/1 values into a string:

df["Binary_Feature"] = df.filter(regex="V\d+").apply(
    lambda row: "".join(row.astype(str)),
    axis=1
)

This produces values like 000100010100001100010 for each row, which can be useful for certain modeling tasks.

Pick the method that best aligns with your downstream work (classification, clustering, etc.)!

内容的提问来源于stack exchange,提问作者Doppler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:35:32