基于R语言DataFrame的学生成绩特定波动事件追踪需求
Hey there! Let's work through how to track those two specific score fluctuation patterns for each student using R, building on the score difference table you already have. I'll break this down into practical, easy-to-follow steps.
1. First, Let's Align on Data Structure (Example)
First, let's assume your score difference table looks something like this (adjust to match your actual data):
# Example score difference data frame score_diffs <- data.frame( student_id = c("John", "John", "John", "John", "Jane", "Jane", "Jane"), date = as.Date(c("2023-01-02", "2023-01-03", "2023-01-04", "2023-01-05", "2023-01-02", "2023-01-03", "2023-01-04")), diff_score = c(7, -6, 3, -2, -8, 9, -2) # Positive = score rise, Negative = score drop )
This table tracks daily score changes for each student (e.g., John's score rose 7 points on 2023-01-02, then dropped 6 points on 2023-01-03).
2. Load Required Packages
We'll use dplyr for grouping and data manipulation, purrr for applying functions per student, and tidyr for nesting/unnesting data:
library(dplyr) library(purrr) library(tidyr)
3. Build a State Machine to Track Fluctuations
We'll create a custom function that uses a simple state machine to track when the two target events occur:
- Neutral: No pending fluctuation to track
- Waiting for Drop: We've seen a score rise of 5+ points, now watching for a 5+ drop
- Waiting for Rise: We've seen a score drop of 5+ points, now watching for a 5+ rise
track_fluctuation_events <- function(student_diffs) { # Initialize state and event tracking variables current_state <- "neutral" event_start_date <- NA events <- data.frame( event_type = character(), start_date = Date(), end_date = Date(), stringsAsFactors = FALSE ) # Iterate through each date's score change for (i in 1:nrow(student_diffs)) { current_diff <- student_diffs$diff_score[i] current_date <- student_diffs$date[i] switch(current_state, "neutral" = { if (current_diff >= 5) { # Trigger: Score rose 5+ -> start waiting for a drop current_state <- "waiting_for_drop" event_start_date <- current_date } else if (current_diff <= -5) { # Trigger: Score dropped 5+ -> start waiting for a rise current_state <- "waiting_for_rise" event_start_date <- current_date } }, "waiting_for_drop" = { if (current_diff <= -5) { # Success: We saw a rise then a drop -> record the event events <- rbind(events, data.frame( event_type = "rise_then_drop", start_date = event_start_date, end_date = current_date, stringsAsFactors = FALSE )) # Reset to neutral after recording current_state <- "neutral" event_start_date <- NA } else if (current_diff >= 5) { # Update start date to the latest rise (prioritize most recent upward movement) event_start_date <- current_date } # Ignore small changes (<5) while waiting }, "waiting_for_rise" = { if (current_diff >= 5) { # Success: We saw a drop then a rise -> record the event events <- rbind(events, data.frame( event_type = "drop_then_rise", start_date = event_start_date, end_date = current_date, stringsAsFactors = FALSE )) # Reset to neutral after recording current_state <- "neutral" event_start_date <- NA } else if (current_diff <= -5) { # Update start date to the latest drop event_start_date <- current_date } # Ignore small changes (<-5) while waiting } ) } # Add event counts and fill in zero counts for students with no events if (nrow(events) == 0) { events <- data.frame( event_type = c("rise_then_drop", "drop_then_rise"), start_date = as.Date(NA), end_date = as.Date(NA), event_count = 0, stringsAsFactors = FALSE ) } else { events <- events %>% group_by(event_type) %>% mutate(event_count = n()) %>% ungroup() } return(events) }
4. Apply the Logic to Each Student
Now we'll group the data by student, apply our tracking function, and unnest the results:
student_fluctuation_summary <- score_diffs %>% arrange(student_id, date) %>% # Critical: Ensure dates are in order group_by(student_id) %>% nest() %>% # Nest data per student mutate(events = map(data, track_fluctuation_events)) %>% # Apply function to each student unnest(events) %>% select(-data) %>% # Remove nested raw data arrange(student_id, event_type) # View the final summary print(student_fluctuation_summary)
For our example data, the output will show John has 1 "rise_then_drop" event, and Jane has 1 "drop_then_rise" event, with their respective start/end dates and counts.
5. Key Notes & Adjustments
- Date Order: Always sort your data by
student_idanddatefirst—this logic relies on chronological order to work correctly. - Edge Cases: If your data has non-consecutive dates, first calculate adjacent date score differences (use
dplyr::lag()to computediff_scoreif you haven't already). - Customization: If you want to exclude small intermediate changes (e.g., don't allow any score movement between the rise/drop), adjust the state machine to reset to neutral on small changes instead of ignoring them.
内容的提问来源于stack exchange,提问作者user98235

