ggplot:如何筛选并展示与均值显著偏离的时间序列
筛选与均值线偏离最大的Top10时间序列方案
嘿,我完全理解你的困扰——当受试者数量涨到50个时,密密麻麻的个体线确实会让图表失去可读性。筛选出偏离均值最显著的Top10序列是个非常合理的解决方案,咱们可以通过计算每个个体序列与均值序列的整体偏离程度来实现,具体步骤如下:
1. 合并均值数据与原始数据
你已经生成了均值数据框df.avg.feeling,首先我们要把它和原始数据df.lines.feeling合并,这样每个个体的每个时间点都能匹配到对应的均值,方便后续计算偏离值:
# 合并均值数据到原始数据集,区分原始值和均值 df_merged <- df.lines.feeling %>% dplyr::left_join(df.avg.feeling, by = "date", suffix = c("", "_avg"))
2. 计算个体序列的偏离度得分
针对时间序列的偏离度,有两种常用的计算方式,你可以根据需求选择:
- 平方误差和(SSE):将每个时间点的
(个体值 - 均值)平方后求和,这种方式会放大较大的偏离值,更适合筛选“显著偏离”的序列 - 绝对误差和(SAE):将每个时间点的
|个体值 - 均值|求和,对极端值的敏感度更低,适合看整体的偏离趋势
这里用SSE来演示,它更符合你“识别显著偏离”的需求:
# 按受试者分组,计算每个个体的总偏离度 df_deviation <- df_merged %>% dplyr::group_by(id) %>% dplyr::mutate(deviation_sq = (feeling - feeling_avg)^2) %>% dplyr::summarise(total_deviation = sum(deviation_sq, na.rm = TRUE)) %>% dplyr::arrange(dplyr::desc(total_deviation)) # 按偏离度从大到小排序
3. 提取Top10偏离最大的序列
从上面的偏离度结果中提取前10个受试者ID,再从原始数据中筛选出这些ID的记录:
# 获取偏离度最高的10个受试者ID top10_ids <- df_deviation %>% dplyr::slice_head(n = 10) %>% dplyr::pull(id) # 生成仅包含Top10序列的数据框 df_top10 <- df.lines.feeling %>% dplyr::filter(id %in% top10_ids)
可选:标准化后的偏离度(Z-score法)
如果不同时间点的均值波动较大,用Z-score计算偏离度会更公平——它把每个时间点的数值转换为相对于该时间点整体样本的偏离程度,再求和:
# 先计算每个时间点的均值和标准差 df_date_stats <- df.lines.feeling %>% dplyr::group_by(date) %>% dplyr::summarise(mean_feeling = mean(feeling), sd_feeling = sd(feeling)) # 合并数据并计算每个点的Z-score,再求和绝对值作为总偏离度 df_z_deviation <- df.lines.feeling %>% dplyr::left_join(df_date_stats, by = "date") %>% dplyr::group_by(id) %>% dplyr::mutate(z_score = (feeling - mean_feeling)/sd_feeling) %>% dplyr::summarise(total_z_deviation = sum(abs(z_score), na.rm = TRUE)) %>% dplyr::arrange(dplyr::desc(total_z_deviation)) # 提取Top10并生成对应数据框 top10_z_ids <- df_z_deviation %>% dplyr::slice_head(n=10) %>% dplyr::pull(id) df_top10_z <- df.lines.feeling %>% dplyr::filter(id %in% top10_z_ids)
绘图验证
现在用筛选后的df_top10绘图,就能清晰看到偏离最大的个体序列和均值线的对比了:
library(ggplot2) ggplot() + # 绘制Top10个体线 geom_line(data = df_top10, aes(x = date, y = feeling, color = factor(id)), alpha = 0.8, linewidth = 1) + # 叠加均值线 geom_line(data = df.avg.feeling, aes(x = date, y = feeling), linewidth = 2, color = "black") + theme_minimal() + labs(color = "Subject ID", title = "Top 10 Deviating Subject Progress vs Mean")
内容的提问来源于stack exchange,提问作者r0berts
相关产品推荐
相关产品推荐

