在R中对指定列子集统计单词数并生成新列
统计文本列单词数并生成总数列的解决方案
问题描述
我有如下数据框:
structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", "This is no text"), y = c("What is text?", "Can it eat text?", "Maybe I will try.")), class = "data.frame", row.names = c(NA, -3L))
需求:统计x和y列的单词数并求和,生成包含每行总单词数的新列z,同时要支持数据子集筛选,且能处理含NA值的列。预期结果如下:
structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", "This is no text"), y = c("What is text?", "Can it eat text?", "Maybe I will try."), z = c("6", "8", "8")), class = "data.frame", row.names = c(NA, -3L))
我尝试过结合正则表达式的str_count(" ")与across或apply,但没得到正确结果。
解决方案
方法1:tidyverse工具链(推荐)
利用stringr的str_count匹配单词,结合dplyr的across批量处理列,同时兼容NA值:
library(dplyr) library(stringr) # 加载原始数据 df <- structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", "This is no text"), y = c("What is text?", "Can it eat text?", "Maybe I will try.")), class = "data.frame", row.names = c(NA, -3L)) # 生成总单词数列z df <- df %>% mutate( # 对x、y列统计单词数:NA值记为0,非NA则匹配所有单词字符计数 across(c(x, y), ~ifelse(is.na(.), 0, str_count(., "\\w+")), .names = "cnt_{.col}"), # 求和并转为字符型(匹配预期结果格式) z = as.character(rowSums(select(., starts_with("cnt")))), # 移除中间计数列 .keep = "unused" ) print(df)
其中\\w+正则表达式会匹配所有字母、数字组成的单词,自动忽略标点符号,确保单词数统计准确。
验证NA值处理
如果数据包含NA,比如修改部分值为NA:
df_with_na <- df %>% mutate(x[2] = NA) df_with_na <- df_with_na %>% mutate( across(c(x, y), ~ifelse(is.na(.), 0, str_count(., "\\w+")), .names = "cnt_{.col}"), z = as.character(rowSums(select(., starts_with("cnt")))), .keep = "unused" ) print(df_with_na)
此时第二行x为NA,统计时按0计算,y列单词数为4,最终z值为"4",符合需求。
子集筛选示例
如果只需处理特定子集(比如g == "1"的行),在处理前添加filter即可:
df_subset <- df %>% filter(g == "1") %>% mutate( across(c(x, y), ~ifelse(is.na(.), 0, str_count(., "\\w+")), .names = "cnt_{.col}"), z = as.character(rowSums(select(., starts_with("cnt")))), .keep = "unused" ) print(df_subset)
方法2:基础R实现
如果不想依赖tidyverse,用基础R的apply和strsplit也能实现:
# 定义单词统计函数,处理NA count_words <- function(text) { if (is.na(text)) return(0) # 按任意空白符分割后计数 length(strsplit(text, "\\s+")[[1]]) } # 加载原始数据 df <- structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", "This is no text"), y = c("What is text?", "Can it eat text?", "Maybe I will try.")), class = "data.frame", row.names = c(NA, -3L)) # 生成z列 df$z <- as.character(apply(df[, c("x", "y")], 1, function(row) sum(sapply(row, count_words)))) print(df)
内容的提问来源于stack exchange,提问作者flxflks
相关产品推荐
相关产品推荐

