使用Pandas计算分数变化并添加新列:个体心理健康分数跨时间点差值计算方案问询
Solution
To calculate the mental health score change between Timepoint 3 and Timepoint 1 for each user (and display this value only in the Timepoint 3 row), here's a straightforward approach using pandas:
Step-by-Step Breakdown
- Isolate Key Timepoints: Extract the mental health scores for each user at Timepoint 1 and Timepoint 3.
- Compute Score Differences: Calculate the difference between each user's Timepoint 3 score and Timepoint 1 score.
- Integrate Changes into Original Data: Merge the computed differences back into your DataFrame, then clear the values for non-Timepoint 3 rows to match your desired output format.
Code Implementation
import pandas as pd # Initialize your sample DataFrame data = { 'User': ['Bill', 'Bill', 'Bill', 'Wiz', 'Wiz', 'Wiz', 'Sam', 'Sam', 'Sam'], 'Timepoint': [1, 2, 3, 1, 2, 3, 1, 2, 3], 'Mental Health Score': [5, 10, 15, 10, 10, 15, 5, 5, 5] } df = pd.DataFrame(data) # Calculate the score change (TP3 - TP1) for each user tp1_scores = df[df['Timepoint'] == 1].set_index('User')['Mental Health Score'] tp3_scores = df[df['Timepoint'] == 3].set_index('User')['Mental Health Score'] change_values = (tp3_scores - tp1_scores).reset_index() change_values.columns = ['User', 'Change in Mental Health (TP1 and 3)'] # Merge the change data into the original DataFrame df = df.merge(change_values, on='User', how='left') # Empty the change column for rows that aren't Timepoint 3 df.loc[df['Timepoint'] != 3, 'Change in Mental Health (TP1 and 3)'] = '' # View the final result print(df)
Final Output
User Timepoint Mental Health Score Change in Mental Health (TP1 and 3) 0 Bill 1 5 1 Bill 2 10 2 Bill 3 15 10 3 Wiz 1 10 4 Wiz 2 10 5 Wiz 3 15 5 6 Sam 1 5 7 Sam 2 5 8 Sam 3 5 0
This method scales well if you add more users to your dataset and ensures the change value only appears in the relevant row, just as you requested.
内容的提问来源于stack exchange,提问作者leahyota
相关产品推荐
相关产品推荐

