You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计数据框中各样本的唯一动物类别并生成对应列表?

解决专属样本动物的统计与列表生成问题

看起来你之前用的sapply(data, function(x) length(unique(x)))其实没抓住核心需求——这个代码只是统计了每个样本列里的不同值的数量(比如0和1,结果肯定都是2),完全不是你要的仅属于该样本的唯一动物数量。下面我给你两种高效的解决方案,不管数据集多大都能稳定运行:

方法一:基础R实现(无需额外包)

先模拟你的数据结构方便演示:

# 模拟你的数据(替换成你实际的data即可)
data <- data.frame(
  Animal = c("Cat", "Dog", "Rabbit", "Mouse", "Snake"),
  Sample1 = c(0, 0, 0, 1, 1),
  Sample2 = c(0, 0, 0, 1, 1),
  Sample3 = c(0, 1, 0, 1, 0),
  Sample4 = c(1, 0, 1, 1, 1),
  stringsAsFactors = FALSE
)

步骤分解:

  1. 找出每个动物对应的所有样本:遍历每一行(动物),提取值为1的样本列名
animal_samples <- apply(data[, -1], 1, function(x) colnames(data[, -1])[x == 1])
  1. 筛选仅属于单个样本的动物:只保留那些只出现在一个样本里的动物
unique_animals <- animal_samples[sapply(animal_samples, length) == 1]
  1. 按样本分组并统计:把这些专属动物分配到对应的样本,统计数量和生成列表
# 分组
sample_unique_animals <- split(names(unique_animals), unlist(unique_animals))
# 统计数量
sample_unique_count <- sapply(sample_unique_animals, length)
  1. 查看结果:
# 打印数量
cat("每个样本的专属唯一动物数量:\n")
print(sample_unique_count)

# 打印列表
cat("\n每个样本的专属唯一动物列表:\n")
for(sample in names(sample_unique_animals)){
  cat(paste0(sample, ": ", paste(sample_unique_animals[[sample]], collapse = ", "), "\n"))
}

运行后你会得到:

每个样本的专属唯一动物数量:
Sample3 Sample4 
      1       2 

每个样本的专属唯一动物列表:
Sample3: Dog
Sample4: Cat, Rabbit

完全符合你描述的预期——Sample1和Sample2没有专属动物,Sample3有Dog,Sample4有Cat和Rabbit。

方法二:Tidyverse实现(更适合大数据集)

如果你的数据集非常大,推荐用tidyverse的函数,代码更简洁且性能优化更好:

library(tidyverse)

result <- data %>%
  # 把宽表转成长表,方便分组处理
  pivot_longer(cols = starts_with("Sample"), names_to = "Sample", values_to = "Present") %>%
  # 只保留存在该动物的样本记录
  filter(Present == 1) %>%
  # 按动物分组,统计每个动物出现在多少个样本里
  group_by(Animal) %>%
  mutate(n_samples = n()) %>%
  # 筛选仅出现在一个样本里的动物
  filter(n_samples == 1) %>%
  # 按样本分组,统计数量并生成动物列表
  group_by(Sample) %>%
  summarise(
    Unique_Animal_Count = n(),
    Unique_Animals = paste(Animal, collapse = ", ")
  )

# 查看结果
print(result)

输出结果是一个数据框,清晰展示每个样本的专属动物数量和列表:

# A tibble: 2 × 3
  Sample  Unique_Animal_Count Unique_Animals
  <chr>                 <int> <chr>         
1 Sample3                   1 Dog           
2 Sample4                   2 Cat, Rabbit   

关键说明

你之前的方法错误在于从列的角度统计唯一值,但我们需要的是从行(动物)的角度判断它是否仅属于某一列(样本),再反向汇总到样本上。上面两种方法都能准确实现你的需求,大数据集下推荐用tidyverse的方案,代码可读性和运行效率都更优。

内容的提问来源于stack exchange,提问作者Hannah9777

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 09:32:33