如何统计R数据框中每行含有效文本值的列数?
R统计每行有效值列数的实现方法
原始数据框(修正缺失值格式)
注:原始代码中的"NA"为字符串,根据说明实际应为缺失值,因此修正后的正确数据框定义如下:
df <- data.frame( id = c(1,2,3), text_en_1 = c("Google", "Yahoo", "Amazon"), text_en_2 = c("Amazon", "Yahoo", NA), text_en_3 = c("Ieekkj", NA, NA), stringsAsFactors = FALSE )
需求说明
基于id列,统计每行中除id外的列里包含有效值(非缺失值)的数量,最终得到包含id和number_of_columns_with_text列的结果数据框,期望输出:
data.frame(id = c(1,2,3), number_of_columns_with_text = c(3,2,1))
解决方案
方法一:基础R实现
利用rowSums()函数直接统计每行非缺失值的数量:
# 计算每行非id列的有效值数量 df$number_of_columns_with_text <- rowSums(!is.na(df[, -1])) # 筛选出目标列 result <- df[, c("id", "number_of_columns_with_text")] print(result)
方法二:dplyr包实现
如果使用tidyverse生态,可通过dplyr实现更简洁的代码:
library(dplyr) result <- df %>% # 统计所有以text_en开头的列的非缺失值数量 mutate(number_of_columns_with_text = rowSums(!is.na(select(., starts_with("text_en"))))) %>% # 保留需要的列 select(id, number_of_columns_with_text) print(result)
输出结果
运行上述任意一种方法,都会得到如下结果:
id number_of_columns_with_text 1 1 3 2 2 2 3 3 1
内容的提问来源于stack exchange,提问作者Erik Brole
相关产品推荐
相关产品推荐

