You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术求助:基于基线绘制均值并执行线性回归的数据预处理方法

数据预处理与可视化解决方案

我们可以用dplyr和tidyr完成数据聚合与基线匹配,用ggplot2实现可视化,以下是完整步骤:

1. 加载工具包并创建初始数据

# 安装并加载所需包(首次使用需执行安装命令)
# install.packages(c("dplyr", "tidyr", "ggplot2"))
library(dplyr)
library(tidyr)
library(ggplot2)

# 创建初始数据框
ID <- c("A","A","A","A","B","B","B","B","C","C","C","C","C")
SampleSection <- c("Base", "First", "Second","Second","Base","First","First","Second","Base","First","First","Second","Second")
lnCort <- c(7.26, 7.68, 7.73, 7.80, 7.95, 7.16, 6.88, 7.81, 7.75, 7.75, 7.40, 8.43, 7.18)
df <- data.frame(ID,SampleSection,lnCort)

2. 计算分组均值并匹配基线

先按ID和SampleSection分组计算均值,再将数据转为宽格式,让每个ID的基线值与处理组均值处于同一行:

processed_df <- df %>%
  # 按ID和分组计算lnCort均值
  group_by(ID, SampleSection) %>%
  summarise(mean_lnCort = mean(lnCort), .groups = "drop") %>%
  # 转换为宽格式,实现基线与处理组的一一匹配
  pivot_wider(names_from = SampleSection, values_from = mean_lnCort,
              names_prefix = "mean_")

处理后的数据结构为:每个ID对应一行,包含mean_Base(基线原始值)、mean_First(First组均值)、mean_Second(Second组均值),完全满足后续分析需求。

3. 绘制指定散点图

图1:Base组值 vs First组均值

ggplot(processed_df, aes(x = mean_Base, y = mean_First)) +
  geom_point(size = 3) +
  labs(x = "Base组lnCort值", y = "First组lnCort均值", title = "Base组与First组均值对比") +
  theme_minimal()

图2:Base组值 vs Second组均值

ggplot(processed_df, aes(x = mean_Base, y = mean_Second)) +
  geom_point(size = 3) +
  labs(x = "Base组lnCort值", y = "Second组lnCort均值", title = "Base组与Second组均值对比") +
  theme_minimal()

补充提示

后续执行线性回归时,直接调用处理后的数据即可,示例代码:

# First组与Base组的线性回归
lm_first <- lm(mean_First ~ mean_Base, data = processed_df)
summary(lm_first)

# Second组与Base组的线性回归
lm_second <- lm(mean_Second ~ mean_Base, data = processed_df)
summary(lm_second)

内容的提问来源于stack exchange,提问作者SmithM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 17:42:15