You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame分组列条件:按指定规则生成新df1的技术需求

Solution to Create df1 Based on Grouping Rules

Got it, let's work through this problem together. Here's a practical way to build your target DataFrame df1 following the grouping logic you specified:

Step-by-Step Approach

The core idea is to split the data by the A column (Red/White), then for each group:

  • Iterate over every unique value in column B
  • First check if there's a row where C matches the current B value — if yes, pick that row
  • If no match for C=B, fall back to picking the row where D matches the current B value

Code Implementation

First, let's start with a sample DataFrame (replace this with your actual data):

import pandas as pd

# Sample input DataFrame
df = pd.DataFrame({
    'A': ['Red', 'Red', 'Red', 'White', 'White', 'White'],
    'B': ['X', 'Y', 'Z', 'X', 'Y', 'Z'],
    'C': ['X', 'Z', 'W', 'Y', 'X', 'Z'],
    'D': ['Y', 'X', 'Z', 'Z', 'Y', 'X'],
    'Value': [10, 20, 30, 40, 50, 60]
})

Next, define a function to process each group, then apply it with groupby:

def process_a_group(group):
    # Get all unique B values in the current group
    unique_b = group['B'].unique()
    selected_rows = []
    
    for b_val in unique_b:
        # Check for C matching B first
        c_match = group[group['C'] == b_val]
        if not c_match.empty:
            # Pick the first matching row (adjust if you need multiple/aggregated rows)
            selected_rows.append(c_match.iloc[0])
        else:
            # Fall back to D matching B
            d_match = group[group['D'] == b_val]
            if not d_match.empty:
                selected_rows.append(d_match.iloc[0])
    
    # Convert selected rows to a DataFrame and reset index
    return pd.DataFrame(selected_rows).reset_index(drop=True)

# Generate df1 by applying the function to each A group
df1 = df.groupby('A').apply(process_a_group).reset_index(drop=True)

# Check the result
print(df1)

Key Notes

  • Handling multiple matches: If your data has multiple rows where C=B (or D=B) for the same B value in a group, the code picks the first one. You can modify this to use iloc[-1] for the last match, or use aggregation functions like groupby().agg() if you need to combine values.
  • Missing matches: If a B value has neither a C nor D match, it won't appear in df1. If you need to handle this case (e.g., fill with NaNs), you can add an extra else block to append a default row.

内容的提问来源于stack exchange,提问作者Tie_24

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:04:12