You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言处理Eurostat数据:地缘政治实体名称缩写问题求助

问题原因与解决方案

未处理值的核心原因

  1. Euro area – 20 countries (from 2023)匹配失败:你的正则表达式中使用的是普通短破折号-,但该字符串里的是长破折号–,二者ASCII码不同,导致正则无法匹配。
  2. European Economic Area (...)匹配失败:startsWith是严格前缀匹配,若字符串存在不可见字符、大小写细微差异(比如开头多余空格),就会匹配失效;相比之下grepl的正则匹配更灵活容错。

修正后的完整代码

df <- df %>% mutate(`Geopolitical entity (reporting)` =
           case_when(
             # 统一Germany的变体写法
             grepl("Germany", `Geopolitical entity (reporting)`, ignore.case = TRUE) ~ "Germany",
             # 处理所有Euro开头的条目:兼容各种破折号、多空格、大小写
             grepl("^Euro", `Geopolitical entity (reporting)`, ignore.case = TRUE) ~
               sub("^([A-Za-z])[A-Za-z]*\\s+([A-Za-z])[A-Za-z]*\\s+[-–—]\\s+(\\d+).*", 
                   "\\U\\1\\U\\2\\3", 
                   `Geopolitical entity (reporting)`, 
                   perl = TRUE),
             # 统一欧洲经济区的缩写为EEA,兼容大小写和前缀匹配
             grepl("^European Economic Area", `Geopolitical entity (reporting)`, ignore.case = TRUE) ~ "EEA",
             # 其余未匹配的原始值保持不变
             TRUE ~ `Geopolitical entity (reporting)`))

关键调整说明

  • 破折号兼容:将\\s-\\s替换为\\s+[-–—]\\s+,其中[-–—]覆盖了短破折号-、长破折号–和全角破折号—;\\s+匹配一个或多个空格,适配Eurostat数据中可能的格式差异。
  • Euro开头匹配优化:用grepl("^Euro", ..., ignore.case = TRUE)替代startsWith,兼容Euro/euro开头的情况,避免大小写导致的匹配遗漏。
  • 欧洲经济区匹配优化:改用grepl做前缀匹配并开启大小写不敏感,解决startsWith的严格匹配限制。

内容的提问来源于stack exchange,提问作者erised

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 09:47:31