You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中对指定列子集统计单词数并生成新列

统计文本列单词数并生成总数列的解决方案

问题描述

我有如下数据框:

structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", 
"This is no text"), y = c("What is text?", "Can it eat text?", 
"Maybe I will try.")), class = "data.frame", row.names = c(NA, 
-3L))

需求:统计x和y列的单词数并求和,生成包含每行总单词数的新列z,同时要支持数据子集筛选,且能处理含NA值的列。预期结果如下:

structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", 
"This is no text"), y = c("What is text?", "Can it eat text?", 
"Maybe I will try."), z = c("6", "8", "8")), class = "data.frame", row.names = c(NA, 
-3L))

我尝试过结合正则表达式的str_count(" ")与across或apply,但没得到正确结果。

解决方案

方法1:tidyverse工具链(推荐)

利用stringr的str_count匹配单词,结合dplyr的across批量处理列,同时兼容NA值:

library(dplyr)
library(stringr)

# 加载原始数据
df <- structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", 
"This is no text"), y = c("What is text?", "Can it eat text?", 
"Maybe I will try.")), class = "data.frame", row.names = c(NA, 
-3L))

# 生成总单词数列z
df <- df %>%
  mutate(
    # 对x、y列统计单词数:NA值记为0,非NA则匹配所有单词字符计数
    across(c(x, y), ~ifelse(is.na(.), 0, str_count(., "\\w+")), .names = "cnt_{.col}"),
    # 求和并转为字符型(匹配预期结果格式)
    z = as.character(rowSums(select(., starts_with("cnt")))),
    # 移除中间计数列
    .keep = "unused"
  )

print(df)

其中\\w+正则表达式会匹配所有字母、数字组成的单词,自动忽略标点符号,确保单词数统计准确。

验证NA值处理

如果数据包含NA,比如修改部分值为NA:

df_with_na <- df %>% mutate(x[2] = NA)

df_with_na <- df_with_na %>%
  mutate(
    across(c(x, y), ~ifelse(is.na(.), 0, str_count(., "\\w+")), .names = "cnt_{.col}"),
    z = as.character(rowSums(select(., starts_with("cnt")))),
    .keep = "unused"
  )

print(df_with_na)

此时第二行x为NA,统计时按0计算,y列单词数为4,最终z值为"4",符合需求。

子集筛选示例

如果只需处理特定子集(比如g == "1"的行),在处理前添加filter即可:

df_subset <- df %>%
  filter(g == "1") %>%
  mutate(
    across(c(x, y), ~ifelse(is.na(.), 0, str_count(., "\\w+")), .names = "cnt_{.col}"),
    z = as.character(rowSums(select(., starts_with("cnt")))),
    .keep = "unused"
  )

print(df_subset)

方法2:基础R实现

如果不想依赖tidyverse,用基础R的apply和strsplit也能实现:

# 定义单词统计函数,处理NA
count_words <- function(text) {
  if (is.na(text)) return(0)
  # 按任意空白符分割后计数
  length(strsplit(text, "\\s+")[[1]])
}

# 加载原始数据
df <- structure(list(g = c("1", "2", "3"), x = c("This is text.", "This is text too.", 
"This is no text"), y = c("What is text?", "Can it eat text?", 
"Maybe I will try.")), class = "data.frame", row.names = c(NA, 
-3L))

# 生成z列
df$z <- as.character(apply(df[, c("x", "y")], 1, function(row) sum(sapply(row, count_words))))

print(df)

内容的提问来源于stack exchange,提问作者flxflks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 22:45:32