如何拆分DataFrame中拼接的grades列并计算参与者得分平均值?
Hey there! Let's solve this problem where you need to compute the average score for each participant from their concatenated grades string column. Here's a straightforward, robust approach using pandas:
Step 1: Example DataFrame Setup
First, let's assume your DataFrame looks something like this (adjust if your actual data has different formatting):
import pandas as pd # Sample data matching your description df = pd.DataFrame({ 'participant': ['a', 'b', 'c'], 'grades': ['4,7,8', '5,9', '3,6,2,7'] })
Step 2: Compute the Average Score
We'll use apply() with a custom function to handle splitting the string, converting to integers, and calculating the average. This version also accounts for edge cases like extra spaces, empty entries, or non-numeric values:
def get_grade_average(grade_string): try: # Split the string, strip whitespace, filter out empty entries, convert to integers scores = [int(score.strip()) for score in grade_string.split(',') if score.strip()] # Calculate average only if there are valid scores return sum(scores) / len(scores) if scores else None except ValueError: # Return None if any entry can't be converted to integer return None # Add the new average column to the DataFrame df['avg_grades'] = df['grades'].apply(get_grade_average)
Step 3: Result
After running the code, your DataFrame will have the new avg_grades column with the calculated averages:
participant grades avg_grades 0 a 4,7,8 6.333333 1 b 5,9 7.000000 2 c 3,6,2,7 4.500000
Simplified Version (No Edge Case Handling)
If you're confident your grades column has no messy data (no extra spaces, no non-numeric values), you can use a shorter lambda function:
df['avg_grades'] = df['grades'].apply( lambda x: sum(int(num) for num in x.split(',')) / len(x.split(',')) )
内容的提问来源于stack exchange,提问作者Rtist

