如何用groupby与筛选DataFrame生成新列,及实现ICU心率数据纳入规则
Hey there! Let's break down your two questions with practical, code-driven solutions since that's the best way to tackle these pandas/data processing tasks.
groupby and Filtering Working with groupby to generate new columns usually boils down to either calculating group-level stats and mapping them back to individual rows, or filtering groups that meet certain criteria and flagging rows accordingly. Here are two common scenarios with pandas examples:
Scenario 1: Add group-level statistics as a new column
Use transform() when you want to compute a statistic (like mean, max, sum) for each group and attach that value to every row in the group. This keeps your original DataFrame structure intact:
import pandas as pd # Sample sales data df = pd.DataFrame({ 'Region': ['North', 'North', 'South', 'South', 'South'], 'Sales': [1200, 1500, 800, 950, 1100] }) # Add a column with the average sales per region df['Region_Avg_Sales'] = df.groupby('Region')['Sales'].transform('mean') # Add a flag for rows where sales exceed the regional average df['Above_Region_Avg'] = df['Sales'] > df['Region_Avg_Sales']
Scenario 2: Flag rows based on group-level filters
If you only want to mark rows that belong to groups meeting a specific condition, use filter() to identify qualifying groups first, then map that back to your DataFrame:
# Find regions where total sales are over 2500 qualified_regions = df.groupby('Region').filter(lambda g: g['Sales'].sum() > 2500)['Region'].unique() # Add a column indicating if the row's region is qualified df['Region_Qualified'] = df['Region'].isin(qualified_regions)
This is a bit more nuanced since we're dealing with time series and specific rules around missing values. Let's assume your dataset has columns like patient_id, timestamp, and heart_rate. Here's a step-by-step solution to implement your requirements:
Step 1: Prep the data
First, always sort your data by patient and timestamp—time series processing depends on ordered data:
import pandas as pd import numpy as np # Sample ICU data (mimicking your dataset structure) data = { 'patient_id': [1,1,1,1,1,2,2,2,2], 'timestamp': pd.to_datetime([ '2024-01-01 10:00', '2024-01-01 10:20', '2024-01-01 10:40', '2024-01-01 11:10', '2024-01-01 11:30', '2024-01-01 08:00', '2024-01-01 08:30', '2024-01-01 09:10', '2024-01-01 09:30' ]), 'heart_rate': [85, 92, 95, np.nan, 91, 88, 90, np.nan, 87] } df = pd.DataFrame(data) # Sort by patient and timestamp df = df.sort_values(['patient_id', 'timestamp']).reset_index(drop=True)
Step 2: Define a function to process each patient's data
We'll write a custom function to check each patient's time series against your inclusion rules, then return only the qualifying rows (starting from the first valid measurement):
def process_patient_records(patient_df): # Mark rows where heart rate is >= 90 patient_df['hr_ge90'] = patient_df['heart_rate'] >= 90 # Get all indices where heart rate first meets the >=90 threshold qualifying_start_indices = patient_df[patient_df['hr_ge90']].index for start_idx in qualifying_start_indices: start_time = patient_df.loc[start_idx, 'timestamp'] one_hour_later = start_time + pd.Timedelta(hours=1) # Check all measurements between start time and 1 hour later hourly_window = patient_df[(patient_df['timestamp'] >= start_time) & (patient_df['timestamp'] <= one_hour_later)] # Ensure all non-missing heart rates in the window are >=90 window_is_valid = hourly_window['heart_rate'].dropna().ge(90).all() if window_is_valid: # Check the first measurement after the 1-hour window post_window_records = patient_df[patient_df['timestamp'] >= one_hour_later] if len(post_window_records) == 0: # No measurements after 1 hour—default to qualifying, return rows from start return patient_df.loc[start_idx:] else: first_post_measurement = post_window_records.iloc[0] # If the first post measurement is missing or >=90, qualify if pd.isna(first_post_measurement['heart_rate']) or first_post_measurement['heart_rate'] >=90: return patient_df.loc[start_idx:] # If no qualifying period found, return empty DataFrame for this patient return pd.DataFrame(columns=patient_df.columns) # Apply the function to each patient group final_filtered_df = df.groupby('patient_id').apply(process_patient_records).reset_index(drop=True)
Key Notes:
- This function checks each potential starting point where heart rate hits >=90, then verifies the 1-hour window and post-window rules.
- If there are no measurements after the 1-hour window, we default to including the patient's records starting from the first valid measurement (per your requirement).
- You can adjust the time delta (e.g.,
pd.Timedelta(minutes=60)is the same as hours=1) if your timestamp uses finer granularity. - If you need to keep the entire patient's record instead of starting from the first qualifying point, just change
return patient_df.loc[start_idx:]toreturn patient_df.
内容的提问来源于stack exchange,提问作者Skum

