You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用groupby与筛选DataFrame生成新列,及实现ICU心率数据纳入规则

Hey there! Let's break down your two questions with practical, code-driven solutions since that's the best way to tackle these pandas/data processing tasks.

1. Creating New Columns with groupby and Filtering

Working with groupby to generate new columns usually boils down to either calculating group-level stats and mapping them back to individual rows, or filtering groups that meet certain criteria and flagging rows accordingly. Here are two common scenarios with pandas examples:

Scenario 1: Add group-level statistics as a new column

Use transform() when you want to compute a statistic (like mean, max, sum) for each group and attach that value to every row in the group. This keeps your original DataFrame structure intact:

import pandas as pd

# Sample sales data
df = pd.DataFrame({
    'Region': ['North', 'North', 'South', 'South', 'South'],
    'Sales': [1200, 1500, 800, 950, 1100]
})

# Add a column with the average sales per region
df['Region_Avg_Sales'] = df.groupby('Region')['Sales'].transform('mean')

# Add a flag for rows where sales exceed the regional average
df['Above_Region_Avg'] = df['Sales'] > df['Region_Avg_Sales']

Scenario 2: Flag rows based on group-level filters

If you only want to mark rows that belong to groups meeting a specific condition, use filter() to identify qualifying groups first, then map that back to your DataFrame:

# Find regions where total sales are over 2500
qualified_regions = df.groupby('Region').filter(lambda g: g['Sales'].sum() > 2500)['Region'].unique()

# Add a column indicating if the row's region is qualified
df['Region_Qualified'] = df['Region'].isin(qualified_regions)
2. Processing ICU Heart Rate Data with Your Inclusion Criteria

This is a bit more nuanced since we're dealing with time series and specific rules around missing values. Let's assume your dataset has columns like patient_id, timestamp, and heart_rate. Here's a step-by-step solution to implement your requirements:

Step 1: Prep the data

First, always sort your data by patient and timestamp—time series processing depends on ordered data:

import pandas as pd
import numpy as np

# Sample ICU data (mimicking your dataset structure)
data = {
    'patient_id': [1,1,1,1,1,2,2,2,2],
    'timestamp': pd.to_datetime([
        '2024-01-01 10:00', '2024-01-01 10:20', '2024-01-01 10:40', 
        '2024-01-01 11:10', '2024-01-01 11:30', '2024-01-01 08:00',
        '2024-01-01 08:30', '2024-01-01 09:10', '2024-01-01 09:30'
    ]),
    'heart_rate': [85, 92, 95, np.nan, 91, 88, 90, np.nan, 87]
}
df = pd.DataFrame(data)

# Sort by patient and timestamp
df = df.sort_values(['patient_id', 'timestamp']).reset_index(drop=True)

Step 2: Define a function to process each patient's data

We'll write a custom function to check each patient's time series against your inclusion rules, then return only the qualifying rows (starting from the first valid measurement):

def process_patient_records(patient_df):
    # Mark rows where heart rate is >= 90
    patient_df['hr_ge90'] = patient_df['heart_rate'] >= 90
    
    # Get all indices where heart rate first meets the >=90 threshold
    qualifying_start_indices = patient_df[patient_df['hr_ge90']].index
    
    for start_idx in qualifying_start_indices:
        start_time = patient_df.loc[start_idx, 'timestamp']
        one_hour_later = start_time + pd.Timedelta(hours=1)
        
        # Check all measurements between start time and 1 hour later
        hourly_window = patient_df[(patient_df['timestamp'] >= start_time) & (patient_df['timestamp'] <= one_hour_later)]
        # Ensure all non-missing heart rates in the window are >=90
        window_is_valid = hourly_window['heart_rate'].dropna().ge(90).all()
        
        if window_is_valid:
            # Check the first measurement after the 1-hour window
            post_window_records = patient_df[patient_df['timestamp'] >= one_hour_later]
            
            if len(post_window_records) == 0:
                # No measurements after 1 hour—default to qualifying, return rows from start
                return patient_df.loc[start_idx:]
            else:
                first_post_measurement = post_window_records.iloc[0]
                # If the first post measurement is missing or >=90, qualify
                if pd.isna(first_post_measurement['heart_rate']) or first_post_measurement['heart_rate'] >=90:
                    return patient_df.loc[start_idx:]
    
    # If no qualifying period found, return empty DataFrame for this patient
    return pd.DataFrame(columns=patient_df.columns)

# Apply the function to each patient group
final_filtered_df = df.groupby('patient_id').apply(process_patient_records).reset_index(drop=True)

Key Notes:

  • This function checks each potential starting point where heart rate hits >=90, then verifies the 1-hour window and post-window rules.
  • If there are no measurements after the 1-hour window, we default to including the patient's records starting from the first valid measurement (per your requirement).
  • You can adjust the time delta (e.g., pd.Timedelta(minutes=60) is the same as hours=1) if your timestamp uses finer granularity.
  • If you need to keep the entire patient's record instead of starting from the first qualifying point, just change return patient_df.loc[start_idx:] to return patient_df.

内容的提问来源于stack exchange,提问作者Skum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:13:06