You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于grepl条件提取多子串并匹配计算累计人口的R语言问题

Solution: Flexible Population Calculation Without Hardcoding

Alright, let's tackle this problem with a flexible, non-hardcoded solution that fits all your requirements. We'll use dplyr and stringr from the tidyverse to dynamically extract countries, match them to df2, and calculate total populations—no hardcoded country or zone values needed.

Step 1: First, Set Up Your Data Correctly

First, let's fix the initial data setup (your original code had variables defined after creating the data frame, which would throw an error):

# Define df1
zone <- c("M", "N", "O")
country_name <- c(
  "The USA, Canada & Mexico are part of North America",
  "Canada like Australia is a Commonwealth member",
  "The UK is still finalizing its exit plans from the EU"
)
df1 <- data.frame(zone, country_name, stringsAsFactors = FALSE)

# Define df2
zonal_region <- c("M", "M", "M", "N", "N", "N", "O", "O", "O")
country <- c("USA", "Canada", "Mexico", "Canada", "Australia", "UK", "Australia", "UK", "Canada")
population <- c(323.4 , 36.29, 127.5, 36.29, 24.13, 65.64, 24.13, 65.64, 36.29)
df2 <- data.frame(zonal_region, country, population, stringsAsFactors = FALSE)

Step 2: Clean Up df2 (Avoid Duplicate Entries)

First, we'll ensure df2 has unique (zonal_region, country) pairs to prevent double-counting populations:

library(dplyr)

df2_clean <- df2 %>%
  distinct(zonal_region, country, .keep_all = TRUE)

Step 3: Dynamic Extraction & Population Calculation

This is the core part—we'll process each row of df1 to:

  1. Pull the list of countries for the row's zone from df2_clean
  2. Extract those countries from the country_name text
  3. Match them to df2_clean to sum their populations
library(stringr)

df3 <- df1 %>%
  rowwise() %>%
  # Get all countries associated with the current zone (no hardcoding!)
  mutate(country_list = list(df2_clean$country[df2_clean$zonal_region == zone])) %>%
  # Extract matching countries from country_name using word boundaries to avoid partial matches
  mutate(matched_countries = list(str_extract_all(
    country_name,
    paste0("\\b", country_list, "\\b", collapse = "|")
  )[[1]])) %>%
  # Sum populations for matched countries in the current zone
  mutate(total_population = sum(
    df2_clean$population[
      df2_clean$zonal_region == zone & df2_clean$country %in% matched_countries
    ]
  )) %>%
  # Remove helper columns
  select(-country_list, -matched_countries) %>%
  ungroup()

Step 4: Verify the Output

If you print df3, you'll get exactly the result you're looking for:

print(df3)
#>   zone                                                                 country_name total_population
#> 1    M                The USA, Canada & Mexico are part of North America           487.19
#> 2    N                    Canada like Australia is a Commonwealth member            60.42
#> 3    O The UK is still finalizing its exit plans from the European Union            65.64

Key Benefits of This Approach

  • No hardcoding: All country and zone values are pulled dynamically from df2, so if your data updates (new zones, new countries), the code adapts automatically.
  • Accurate matching: We use \\b (word boundaries) in the regex to ensure we only match full country names (e.g., we won't accidentally extract "UK" from "EU").
  • Clean, readable: The tidyverse syntax makes the logic easy to follow and modify if needed.

内容的提问来源于stack exchange,提问作者Runner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:10:38