You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Z-score计算的箱线图填充:运动员体能数据对比绘图需求

Solution: Z-Score-Based Box Plot with Fill for Athlete Performance Comparison

Let's walk through how to create a box plot that compares your current athlete data to the quartertwo2017 historical dataset, using Z-scores to drive the fill color. First, a quick note: your provided historical data has an incomplete final row (Player 1 5)—I'll exclude that in the example below since it's missing key metrics.

Step 1: Set Up Your Environment & Data

First, load the necessary packages and clean up the historical dataset:

# Load required packages
library(tidyverse)

# Create the cleaned historical dataset (removing the incomplete final row)
quartertwo2017 <- tibble(
  Player.Name = rep("Player 1", 7),
  Date = c("10/9/17", "15/7/17", "2/7/17", "22/7/17", "24/6/17", "29/7/17", "3/9/17"),
  Distance = c(7060.621, 4978.625, 6787.667, 6881.126, 5802.060, 7035.075, 7016.175),
  HIR = c(2506.20, 1596.19, 2048.61, 2065.80, 2204.87, 2085.32, 2659.18),
  V6 = c(12.50, 44.26, 39.67, 31.48, 65.48, 22.56, 66.14)
) %>%
  # Convert Date to proper date format
  mutate(Date = dmy(Date))

# Example current data (replace this with your actual current performance data)
current_data <- tibble(
  Player.Name = "Player 1",
  Date = today(), # Or your actual current date
  Distance = 7200, # Example current value
  HIR = 2400, # Example current value
  V6 = 35 # Example current value
)

Step 2: Calculate Z-Scores for All Metrics

Z-scores let us standardize the data so we can compare performance relative to the historical average. We'll compute Z-scores for both historical and current data:

# Calculate mean and sd for each metric from historical data
historical_summary <- quartertwo2017 %>%
  summarise(
    across(c(Distance, HIR, V6), list(mean = mean, sd = sd), .names = "{.col}_{.fn}")
  )

# Combine historical and current data, then compute Z-scores
combined_data <- bind_rows(quartertwo2017, current_data) %>%
  mutate(
    # Z-score formula: (value - historical mean) / historical sd
    Distance_Z = (Distance - historical_summary$Distance_mean) / historical_summary$Distance_sd,
    HIR_Z = (HIR - historical_summary$HIR_mean) / historical_summary$HIR_sd,
    V6_Z = (V6 - historical_summary$V6_mean) / historical_summary$V6_sd,
    # Label current data for highlighting
    Data_Type = ifelse(Date == max(Date), "Current", "Historical")
  ) %>%
  # Reshape data to long format for easier plotting
  pivot_longer(cols = ends_with("_Z"), names_to = "Metric", values_to = "Z_Score") %>%
  mutate(Metric = str_remove(Metric, "_Z"))

Step 3: Create the Z-Score Box Plot with Fill

Now we'll use ggplot2 to make a box plot where the fill color corresponds to how the current data compares to the historical distribution (via Z-score):

ggplot(combined_data, aes(x = Metric, y = Z_Score)) +
  # Box plot for historical data
  geom_boxplot(fill = "#f0f0f0", alpha = 0.7, width = 0.6) +
  # Point for current data, sized larger and colored by Z-score
  geom_point(
    data = filter(combined_data, Data_Type == "Current"),
    aes(color = Z_Score),
    size = 5,
    shape = 19
  ) +
  # Add a horizontal line at Z=0 (historical average)
  geom_hline(yintercept = 0, linetype = "dashed", color = "red", size = 0.8) +
  # Color scale for Z-score fill (adjust colors as needed)
  scale_color_gradient2(
    midpoint = 0,
    low = "blue", # Below average
    mid = "white", # Average
    high = "darkorange", # Above average
    name = "Z-Score"
  ) +
  # Custom labels and theme
  labs(
    title = "Player 1: Current vs. Q2 2017 Performance (Z-Score)",
    y = "Z-Score (Relative to Q2 2017 Average)",
    x = "Performance Metric",
    caption = "Historical data: quartertwo2017 | Current data: recent performance"
  ) +
  theme_minimal() +
  theme(
    plot.title = element_text(hjust = 0.5, size = 14, face = "bold"),
    axis.title = element_text(size = 12),
    legend.position = "right"
  )

Key Notes:

  • The box plot shows the distribution of historical Z-scores (since we standardized using historical mean/sd)
  • The colored point represents the current performance's Z-score—blue means below historical average, orange means above
  • The dashed red line is at Z=0, the historical average
  • Replace the current_data values with your actual current athlete metrics to get accurate results

内容的提问来源于stack exchange,提问作者user9460505

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:23:01