You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中实现people数据集bio字段与location数据集城市的动态匹配?

动态匹配文本中的城市并生成结果数据集

问题描述

现有两个数据集:

  • people 数据集包含用户ID和个人简介:
people <- structure(list(userID = c(175890530, 178691082, 40228319, 472555502, 
                                  1063565418, 242983504, 3253221155), bio = c("Living in Atlanta", 
                                                                              "Born in Seattle, resident of Phoenix", "Columbus, Ohio", "Bronx born and raised", 
                                                                              "What's up Chicago?!?!", "Product of Los Angeles, taxpayer in St. Louis", 
                                                                              "Go Dallas Cowboys!")), class = "data.frame", row.names = c(NA, 
                                                                                                                                          -7L))
  • location 数据集包含城市及对应州:
location <- structure(list(city = c("Atlanta", "Seattle", "Phoenix", "Columbus", 
                                  "Bronx", "Chicago", "Los Angeles", "St. Louis", "Dallas"), state = c("GA", 
                                                                                                       "WA", "AZ", "OH", "NY", "IL", "CA", "MO", "TX")), class = "data.frame", row.names = c(NA, 
                                                                                                                                                                                             -9L))

需要从people$bio的每一行文本中匹配location$city中的所有城市,生成包含userID、bio和新增city_return字段的complete数据集,要求不能硬编码城市列表(因实际场景中城市数量庞大),最终输出结构如下:

complete <- structure(list(userID = c(175890530, 178691082, 40228319, 472555502, 
1063565418, 242983504, 3253221155), bio = c("Living in Atlanta", 
"Born in Seattle, resident of Phoenix", "Columbus, Ohio", "Bronx born and raised", 
"What's up Chicago?!?!", "Product of Los Angeles, taxpayer in St. Louis", 
"Go Dallas Cowboys!"), city_return = c("Atlanta", "Seattle, Phoenix", 
"Columbus", "Bronx", "Chicago", "Los Angeles, St. Louis", "Dallas"
)), class = "data.frame", row.names = c(NA, -7L))

解决方案

使用dplyr和stringr包实现动态匹配,无需硬编码城市列表:

# 加载所需包
library(dplyr)
library(stringr)

# 动态生成城市匹配的正则模式(按城市名长度降序排列,避免短名误匹配长名前缀)
city_pattern <- location$city %>%
  str_sort(decreasing = TRUE, by = str_length) %>%
  str_c(collapse = "|")

# 处理数据集,提取匹配的城市并合并为逗号分隔的字符串
complete <- people %>%
  mutate(
    city_return = str_extract_all(bio, regex(city_pattern, ignore_case = FALSE)) %>%
      map_chr(~str_c(.x, collapse = ", "))
  )

# 查看结果
complete

关键说明

  • 按城市名长度降序排列正则模式:确保长城市名(如"Los Angeles")优先匹配,避免短名(若存在)错误匹配长名的前缀部分。
  • str_extract_all提取所有匹配的城市,返回列表格式后,通过map_chr结合str_c将列表元素合并为逗号分隔的字符串。
  • 若需要忽略大小写匹配,可将regex中的ignore_case参数改为TRUE。

内容的提问来源于stack exchange,提问作者wizkids121

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 06:15:20