You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言的基因操作:生成频率列及获取次高频等位基因

解决方案:碱基频率计算与次高频等位基因提取

1. 生成各碱基频率列

你之前的宽表转换没问题,但频率计算需要按patient*gene分组,计算每组内的碱基占比。以下是两种实现方式:

方式一:基于已转换的宽表计算

clean_df_mut_counts_wide <- clean_df_mut_counts_wide %>%
  group_by(patient, gene) %>%
  mutate(
    total_count = A + T + C + G,  # 计算每组总突变数
    A_freq = A / total_count,
    T_freq = T / total_count,
    C_freq = C / total_count,
    G_freq = G / total_count
  ) %>%
  ungroup()

方式二:从原始长表直接生成(更灵活)

如果原始长表clean_df_mut_counts包含patient、gene、base、count字段,可以一步到位生成带频率的宽表:

clean_df_with_freq <- clean_df_mut_counts %>%
  group_by(patient, gene) %>%
  mutate(total_count = sum(count), freq = count / total_count) %>%
  pivot_wider(
    names_from = base,
    values_from = c(count, freq),
    names_glue = "{base}_{.value}"  # 自动生成A_count、A_freq这类规范列名
  ) %>%
  ungroup()

2. 获取每个patient*gene分组下的第二高频率等位基因

核心思路是把宽表转成长表,按分组排序后提取次高频记录:

方法一:基于已生成频率列的宽表

second_highest_allele <- clean_df_mut_counts_wide %>%
  # 将频率列转成长格式,保留分组信息
  pivot_longer(
    cols = ends_with("_freq"),
    names_to = "base",
    values_to = "freq",
    names_transform = list(base = ~ gsub("_freq", "", .))  # 去掉碱基名后缀
  ) %>%
  group_by(patient, gene) %>%
  arrange(desc(freq), base) %>%  # 按频率降序排序,同频时按碱基名排序(可选)
  slice(2) %>%  # 取每组第2条记录(即次高频)
  select(patient, gene, second_highest_base = base, second_highest_freq = freq) %>%
  ungroup()

方法二:直接从原始长表处理(更简洁)

second_highest_allele <- clean_df_mut_counts %>%
  group_by(patient, gene) %>%
  mutate(freq = count / sum(count)) %>%  # 直接计算频率
  arrange(desc(freq), base) %>%
  slice(2) %>%
  select(patient, gene, second_highest_base = base, second_highest_freq = freq) %>%
  ungroup()

你之前代码的问题说明

你之前的代码错误地将A/T/C/G的数值作为分组依据,完全偏离了patient*gene的分组需求;同时summarise(n = n())只是统计相同碱基计数组合的行数,和频率计算、次高频提取毫无关系。

内容的提问来源于stack exchange,提问作者analog_kid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 19:06:28