You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用gsub拆分含多词的城市与国家字符串问题求助

解决方案

你的问题出在原正则仅匹配连续字母,没法处理带空格、连字符的多词内容,且没抓住「下划线是唯一分隔符」的核心逻辑——应该把下划线前的所有内容作为城市,下划线后的所有内容作为国家。

方法1:用sub优化正则

sub比gsub更适合此场景,因为每个字符串只有一个分隔下划线,仅需替换一次:

# 示例数据
locations <- c("paris_france", "miami_united states", "new york_united states", "aix-en-provence_france", "auckland_new_zealand")

# 提取城市:匹配开头到第一个下划线的所有内容
city <- sub("^(.*?)_.*$", "\\1", locations)

# 提取国家:匹配最后一个下划线到结尾的所有内容
country <- sub("^.*_(.*)$", "\\1", locations)

方法2:用strsplit分割(更直观)

直接按下划线分割成两部分,分别取第一、第二个元素,无需正则,更不易出错:

# 分割字符串,fixed=TRUE表示下划线是固定字符,而非正则元字符
split_loc <- strsplit(locations, "_", fixed = TRUE)

city <- sapply(split_loc, `[`, 1)
country <- sapply(split_loc, `[`, 2)

验证结果

运行后得到的结果完全符合需求:

  • 城市:"paris" "miami" "new york" "aix-en-provence" "auckland"
  • 国家:"france" "united states" "united states" "france" "new_zealand"

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 22:52:43