如何用R或bash统计不同ID列中各字符串的出现频次
R/Bash实现频数统计转换操作
需求说明
给定如下原始数据表:
ID1 ID2 ID3 ------------- a a b a b b b b b c c c c c d c d d d e d e e
需要转换为按样本值统计各ID列出现次数的结果表:
Samples ID1 ID2 ID3 ------------------- a 2 1 0 b 1 2 3 c 3 2 1 d 2 1 2 e 1 0 2
R实现代码
# 读入原始数据,自动填充空值为NA df <- read.table(text = "ID1 ID2 ID3 a a b a b b b b b c c c c c d c d d d e d e e", header = TRUE, fill = TRUE, na.strings = "") # 提取所有出现过的样本值 all_samples <- unique(na.omit(c(df$ID1, df$ID2, df$ID3))) # 统计各列每个值的出现频次 count_id1 <- table(df$ID1) count_id2 <- table(df$ID2) count_id3 <- table(df$ID3) # 构建结果表,缺失值替换为0 result <- data.frame( Samples = all_samples, ID1 = as.numeric(count_id1[all_samples]), ID2 = as.numeric(count_id2[all_samples]), ID3 = as.numeric(count_id3[all_samples]) ) result[is.na(result)] <- 0 result <- result[order(result$Samples), ] # 输出结果 print(result, row.names = FALSE)
Bash实现代码
先将原始数据保存为input.txt,再执行以下脚本:
# 分别统计三列的各值频次,存入临时文件 sed '1,2d' input.txt | awk '{print $1}' | grep -v '^$' | sort | uniq -c > id1_cnt.txt sed '1,2d' input.txt | awk '{print $2}' | grep -v '^$' | sort | uniq -c > id2_cnt.txt sed '1,2d' input.txt | awk '{print $3}' | grep -v '^$' | sort | uniq -c > id3_cnt.txt # 输出表头 echo -e "Samples\tID1\tID2\tID3" # 遍历所有样本值,匹配各列频次,无匹配则补0 awk '{print $2}' id1_cnt.txt id2_cnt.txt id3_cnt.txt | sort | uniq | while read sample; do id1=$(grep -w "$sample" id1_cnt.txt | awk '{print $1}' || echo 0) id2=$(grep -w "$sample" id2_cnt.txt | awk '{print $1}' || echo 0) id3=$(grep -w "$sample" id3_cnt.txt | awk '{print $1}' || echo 0) echo -e "$sample\t$id1\t$id2\t$id3" done # 清理临时文件 rm -f id1_cnt.txt id2_cnt.txt id3_cnt.txt
内容的提问来源于stack exchange,提问作者Anisur Rahman
相关产品推荐
相关产品推荐

