You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写for循环迭代DataFrame并按条件生成包含指定列的子集DataFrame

Hey there! Let's work through how to create those subset DataFrames you need with a for loop. Here's a straightforward approach that hits all your requirements:

Solution: Iterate Over Sample Columns to Generate Target Subsets

First, let's set up some example data

To make this concrete, let's create a sample DataFrame that matches your use case (with id, reference, multiple sample columns, and some extra columns to ignore):

import pandas as pd

# Example DataFrame to test our logic
df = pd.DataFrame({
    'id': [1, 2, 3, 4, 5],
    'reference': ['ref_A', 'ref_B', 'ref_C', 'ref_D', 'ref_E'],
    'sample 1': [0, 1, 0, 1, 0],
    'sample 2': [1, 0, 1, 0, 1],
    'non_sample_col': ['x', 'y', 'z', 'x', 'y']  # This column won't be included in subsets
})

Step 1: Identify all "sample" columns

First, we need to grab every column in the DataFrame that's named with "sample" (adjust this if your naming pattern is slightly different, like all columns containing "sample" instead of starting with it):

# Get columns that start with "sample" (use str.contains('sample') if needed)
sample_columns = [col for col in df.columns if col.startswith('sample')]

Step 2: Loop through each sample column to create subsets

We'll use a dictionary to store all our generated subsets (this makes it easy to access them later by name). For each sample column, we'll create two subsets: one where the column value is 0, and another where it's 1, each only keeping the id, reference, and current sample column.

# Dictionary to store all our subset DataFrames
subset_dataframes = {}

for col in sample_columns:
    # Subset where current sample column is 0, keep only specified columns
    subset_0 = df[df[col] == 0][['id', 'reference', col]]
    subset_dataframes[f"{col}_value_0"] = subset_0
    
    # Subset where current sample column is 1, keep only specified columns
    subset_1 = df[df[col] == 1][['id', 'reference', col]]
    subset_dataframes[f"{col}_value_1"] = subset_1

Step 3: Check the results

To confirm everything works as expected, you can print out any of the subsets:

# Example: Print the subset for "sample 1" where value is 0
print("Subset for 'sample 1' (value = 0):")
print(subset_dataframes['sample 1_value_0'])

# Example: Print the subset for "sample 2" where value is 1
print("\nSubset for 'sample 2' (value = 1):")
print(subset_dataframes['sample 2_value_1'])

Bonus Tips

  • If your "sample" columns don't start with "sample" (e.g., they're named test_sample_3), replace startswith('sample') with 'sample' in col or col.str.contains('sample') (note: use col.str only if all column names are strings, which they usually are).
  • If you want to save each subset directly to a CSV file instead of storing them in a dictionary, add this line inside the loop:
    subset_0.to_csv(f"{col}_value_0.csv", index=False)
    subset_1.to_csv(f"{col}_value_1.csv", index=False)
    
  • If your sample columns have values other than 0/1 and you need to handle those, you can adjust the filter condition (e.g., df[col] == 'active' for string values).

内容的提问来源于stack exchange,提问作者Aruna Saraswathy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 16:22:45