如何按LabelId=1拆分DataFrame为两个可区分的子集?
看起来你是想把LabelId=1的行拆成两个独立的DataFrame对吧?直接用布尔索引只会把所有符合条件的行捞出来,确实没法直接区分两个子集。这里给你几种常用的解决思路,根据你的实际需求选就行:
1. 按行位置拆分(比如前半部分/后半部分)
如果只是想简单把符合条件的行分成前后两段,先筛选出目标行再拆分:
# 第一步:先提取所有LabelId=1的行 df_label1 = DF_input[DF_input["LabelId"] == 1] # 第二步:按行数拆分,比如取前一半给DF_output1,剩下的给DF_output2 split_point = len(df_label1) // 2 DF_output1 = df_label1.iloc[:split_point] DF_output2 = df_label1.iloc[split_point:]
2. 按额外列的条件拆分
如果你的DataFrame里有其他可以用来区分的列(比如状态、类别、时间戳等),可以在LabelId=1的基础上叠加条件:
# 举个例子:按"Status"列的取值拆分 DF_output1 = DF_input[(DF_input["LabelId"] == 1) & (DF_input["Status"] == "Active")] DF_output2 = DF_input[(DF_input["LabelId"] == 1) & (DF_input["Status"] == "Inactive")]
3. 随机拆分(适合做训练/测试集)
如果需要随机将LabelId=1的行分成两个子集(比如做模型训练和验证),可以用scikit-learn的train_test_split:
from sklearn.model_selection import train_test_split # 先筛选出目标行 df_label1 = DF_input[DF_input["LabelId"] == 1] # 拆分:test_size控制比例,random_state保证结果可复现 DF_output1, DF_output2 = train_test_split(df_label1, test_size=0.3, random_state=42)
4. 按自定义规则拆分(比如指定特定ID)
如果有明确的拆分规则(比如某些特定的记录ID),可以直接用索引匹配:
df_label1 = DF_input[DF_input["LabelId"] == 1] # 假设要把ID在[1001, 1003, 1005]的分到DF_output1,其余到DF_output2 target_ids = [1001, 1003, 1005] DF_output1 = df_label1[df_label1["Id"].isin(target_ids)] DF_output2 = df_label1[~df_label1["Id"].isin(target_ids)]
内容的提问来源于stack exchange,提问作者Carlo Allocca
相关产品推荐
相关产品推荐

