You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ggplot:如何筛选并展示与均值显著偏离的时间序列

筛选与均值线偏离最大的Top10时间序列方案

嘿,我完全理解你的困扰——当受试者数量涨到50个时,密密麻麻的个体线确实会让图表失去可读性。筛选出偏离均值最显著的Top10序列是个非常合理的解决方案,咱们可以通过计算每个个体序列与均值序列的整体偏离程度来实现,具体步骤如下:

1. 合并均值数据与原始数据

你已经生成了均值数据框df.avg.feeling,首先我们要把它和原始数据df.lines.feeling合并,这样每个个体的每个时间点都能匹配到对应的均值,方便后续计算偏离值:

# 合并均值数据到原始数据集,区分原始值和均值
df_merged <- df.lines.feeling %>%
  dplyr::left_join(df.avg.feeling, by = "date", suffix = c("", "_avg"))

2. 计算个体序列的偏离度得分

针对时间序列的偏离度,有两种常用的计算方式,你可以根据需求选择:

  • 平方误差和(SSE):将每个时间点的(个体值 - 均值)平方后求和,这种方式会放大较大的偏离值,更适合筛选“显著偏离”的序列
  • 绝对误差和(SAE):将每个时间点的|个体值 - 均值|求和,对极端值的敏感度更低,适合看整体的偏离趋势

这里用SSE来演示,它更符合你“识别显著偏离”的需求:

# 按受试者分组,计算每个个体的总偏离度
df_deviation <- df_merged %>%
  dplyr::group_by(id) %>%
  dplyr::mutate(deviation_sq = (feeling - feeling_avg)^2) %>%
  dplyr::summarise(total_deviation = sum(deviation_sq, na.rm = TRUE)) %>%
  dplyr::arrange(dplyr::desc(total_deviation)) # 按偏离度从大到小排序

3. 提取Top10偏离最大的序列

从上面的偏离度结果中提取前10个受试者ID,再从原始数据中筛选出这些ID的记录:

# 获取偏离度最高的10个受试者ID
top10_ids <- df_deviation %>%
  dplyr::slice_head(n = 10) %>%
  dplyr::pull(id)

# 生成仅包含Top10序列的数据框
df_top10 <- df.lines.feeling %>%
  dplyr::filter(id %in% top10_ids)

可选:标准化后的偏离度(Z-score法)

如果不同时间点的均值波动较大,用Z-score计算偏离度会更公平——它把每个时间点的数值转换为相对于该时间点整体样本的偏离程度,再求和:

# 先计算每个时间点的均值和标准差
df_date_stats <- df.lines.feeling %>%
  dplyr::group_by(date) %>%
  dplyr::summarise(mean_feeling = mean(feeling),
                   sd_feeling = sd(feeling))

# 合并数据并计算每个点的Z-score,再求和绝对值作为总偏离度
df_z_deviation <- df.lines.feeling %>%
  dplyr::left_join(df_date_stats, by = "date") %>%
  dplyr::group_by(id) %>%
  dplyr::mutate(z_score = (feeling - mean_feeling)/sd_feeling) %>%
  dplyr::summarise(total_z_deviation = sum(abs(z_score), na.rm = TRUE)) %>%
  dplyr::arrange(dplyr::desc(total_z_deviation))

# 提取Top10并生成对应数据框
top10_z_ids <- df_z_deviation %>% dplyr::slice_head(n=10) %>% dplyr::pull(id)
df_top10_z <- df.lines.feeling %>% dplyr::filter(id %in% top10_z_ids)

绘图验证

现在用筛选后的df_top10绘图,就能清晰看到偏离最大的个体序列和均值线的对比了:

library(ggplot2)

ggplot() +
  # 绘制Top10个体线
  geom_line(data = df_top10, aes(x = date, y = feeling, color = factor(id)), alpha = 0.8, linewidth = 1) +
  # 叠加均值线
  geom_line(data = df.avg.feeling, aes(x = date, y = feeling), linewidth = 2, color = "black") +
  theme_minimal() +
  labs(color = "Subject ID", title = "Top 10 Deviating Subject Progress vs Mean")

内容的提问来源于stack exchange,提问作者r0berts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 12:13:12