R语言dplyr按州匹配词表的文本计数问题排查与实现
问题排查与解决方案
错误原因
报错object 'state_abbrev' not found大概率是两个原因:
- 代码里误将数据框的列名
state_abbrevs写成了state_abbrev(少了末尾的s) - 关联词表列表
terms时,没有正确将列表标签与mydf的state_abbrevs字段对应起来
正确实现步骤
我们可以通过将词表列表转换为结构化数据框,再与mydf关联的方式实现需求,具体代码如下(基于tidyverse工具链):
library(tidyverse) # 1. 将terms列表转换为数据框:包含州缩写和对应词项 terms_df <- enframe(terms, name = "state_abbrev", value = "term") %>% unnest(term) # 展开每个州的词项 # 2. 关联mydf和词表,仅保留州缩写匹配的行,判断text是否包含对应词项 result <- mydf %>% inner_join(terms_df, by = c("state_abbrevs" = "state_abbrev")) %>% # 按州缩写匹配 mutate(contains_term = str_detect(text, regex(term, ignore_case = TRUE))) %>% # 忽略大小写匹配 group_by(last_name, first_name, tw_year, tw_month) %>% summarise( total_records = n_distinct(row_number()), # 统计分组内的总记录数 matched_records = sum(contains_term), # 统计匹配到词项的记录数 .groups = "drop" )
代码说明
enframe()把列表转换成两列的数据框,name对应列表的州缩写标签,value对应词项inner_join()确保只保留州缩写匹配的行,过滤掉不匹配的记录str_detect()结合regex(ignore_case = TRUE)实现不区分大小写的字符串匹配,避免漏判- 分组统计时用
n_distinct(row_number())确保每条原始记录只被计数一次(因为关联词表后可能会有重复行)
示例输出
| last_name | first_name | tw_year | tw_month | total_records | matched_records |
|---|---|---|---|---|---|
| Smith | John | 2023 | 10 | 15 | 7 |
| Doe | Jane | 2023 | 10 | 22 | 11 |
内容的提问来源于stack exchange,提问作者Stephanie Ramsey Davis
相关产品推荐
相关产品推荐

