You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr::filter从另一数据框筛选行无数据返回的问题

问题

我尝试使用dplyr::filter函数进行数据筛选:现有存储20个基因名称的数据框top20_res,以及包含多列(其中一列是基因名列ext_gene)的数据框gathered_group1_norm。我希望从gathered_group1_norm中筛选出存在于top20_res中的20个基因及其对应列数据,于是编写了如下代码:

df <- gathered_group1_norm %>% dplyr::filter(top20_res %in% gathered_group1_norm$ext_gene)

但运行后生成的df仅有列名,无实际数据,请问该如何解决?

解决方案

你的代码存在两个核心问题:

  • 逻辑判断顺序颠倒:%in%的正确用法是「待检查值 %in% 目标集合」,你现在写反了,应该检查ext_gene是否属于top20_res的基因集合,而非反过来。
  • 直接使用数据框匹配错误:top20_res是数据框,不能直接用于%in%匹配,需要提取其中存储基因名称的列。

正确代码示例

假设top20_res中存储基因名称的列名为gene(如果你的列名不同,替换成实际列名即可):

# 提取top20_res的基因列,筛选gathered_group1_norm中匹配的行
df <- gathered_group1_norm %>% 
  dplyr::filter(ext_gene %in% top20_res$gene)

如果top20_res只有一列(即基因名称列),可以用dplyr::pull()直接提取向量:

df <- gathered_group1_norm %>% 
  dplyr::filter(ext_gene %in% dplyr::pull(top20_res))

额外验证步骤

如果运行正确代码后仍无数据,需检查:

  • 两个数据框中的基因名称格式是否完全一致:比如大小写、是否包含空格/特殊字符,GeneA和genea会被判定为不同值。
  • 确认两个数据集存在交集:运行以下代码查看共同基因:
intersect(top20_res$gene, gathered_group1_norm$ext_gene)

如果返回空向量,说明两个数据集没有匹配的基因,需检查数据来源或预处理步骤。

内容的提问来源于stack exchange,提问作者Rob Staruch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 21:55:22