You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按LabelId=1拆分DataFrame为两个可区分的子集?

看起来你是想把LabelId=1的行拆成两个独立的DataFrame对吧?直接用布尔索引只会把所有符合条件的行捞出来,确实没法直接区分两个子集。这里给你几种常用的解决思路,根据你的实际需求选就行:

1. 按行位置拆分(比如前半部分/后半部分)

如果只是想简单把符合条件的行分成前后两段,先筛选出目标行再拆分:

# 第一步:先提取所有LabelId=1的行
df_label1 = DF_input[DF_input["LabelId"] == 1]

# 第二步:按行数拆分,比如取前一半给DF_output1,剩下的给DF_output2
split_point = len(df_label1) // 2
DF_output1 = df_label1.iloc[:split_point]
DF_output2 = df_label1.iloc[split_point:]

2. 按额外列的条件拆分

如果你的DataFrame里有其他可以用来区分的列(比如状态、类别、时间戳等),可以在LabelId=1的基础上叠加条件:

# 举个例子:按"Status"列的取值拆分
DF_output1 = DF_input[(DF_input["LabelId"] == 1) & (DF_input["Status"] == "Active")]
DF_output2 = DF_input[(DF_input["LabelId"] == 1) & (DF_input["Status"] == "Inactive")]

3. 随机拆分(适合做训练/测试集)

如果需要随机将LabelId=1的行分成两个子集(比如做模型训练和验证),可以用scikit-learn的train_test_split:

from sklearn.model_selection import train_test_split

# 先筛选出目标行
df_label1 = DF_input[DF_input["LabelId"] == 1]
# 拆分:test_size控制比例,random_state保证结果可复现
DF_output1, DF_output2 = train_test_split(df_label1, test_size=0.3, random_state=42)

4. 按自定义规则拆分(比如指定特定ID)

如果有明确的拆分规则(比如某些特定的记录ID),可以直接用索引匹配:

df_label1 = DF_input[DF_input["LabelId"] == 1]
# 假设要把ID在[1001, 1003, 1005]的分到DF_output1,其余到DF_output2
target_ids = [1001, 1003, 1005]
DF_output1 = df_label1[df_label1["Id"].isin(target_ids)]
DF_output2 = df_label1[~df_label1["Id"].isin(target_ids)]

内容的提问来源于stack exchange,提问作者Carlo Allocca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:16:23