You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按非数字(不含小数点)字符拆分字符串列处理DMS坐标

R实现乱分隔符DMS坐标拆分方案

核心逻辑

你需要的拆分规则可以直接用正则表达式[^0-9.]实现:该正则会匹配所有非数字且不等于小数点的字符,完全符合你的分隔符特征。

实现代码

首先加载必要依赖包:

library(tidyverse)

你的示例数据如下:

df <- tibble(lat=c("22ª29'56.06\"","22°29`53.14\"","22º30'00.43\"","105'29'27.17\""),
             long=c("105º21'37.27\"","105°21'29.48\"","105°21'37.46\''","105°21'39.68"))

方法1:使用tidyr::separate_wider_regex直接拆分

直接通过正则按位置提取度、分、秒数值,不受混乱分隔符干扰,稳定性高:

df_result <- df %>%
  # 拆分纬度列
  separate_wider_regex(
    cols = lat,
    patterns = c(lat_deg = "\\d+", "[^0-9.]+", lat_min = "\\d+", "[^0-9.]+", lat_sec = "[\\d.]+", ".*"),
    cols_remove = FALSE
  ) %>%
  # 拆分经度列
  separate_wider_regex(
    cols = long,
    patterns = c(long_deg = "\\d+", "[^0-9.]+", long_min = "\\d+", "[^0-9.]+", long_sec = "[\\d.]+", ".*"),
    cols_remove = FALSE
  ) %>%
  # 转换为数值类型方便后续坐标转换
  mutate(across(c(lat_deg, lat_min, lat_sec, long_deg, long_min, long_sec), as.numeric))

方法2:使用stringr::str_extract_all通用提取

如果需要更灵活的自定义处理逻辑,可以用批量提取数值的方式实现:

# 自定义提取DMS数值的函数
extract_dms <- function(x) {
  str_extract_all(x, "\\d+\\.?\\d*")[[1]] %>% set_names(c("deg", "min", "sec"))
}

df_result <- df %>%
  rowwise() %>%
  mutate(lat_dms = list(extract_dms(lat)),
         long_dms = list(extract_dms(long))) %>%
  unnest_wider(c(lat_dms, long_dms), names_sep = "_") %>%
  mutate(across(ends_with(c("_deg", "_min", "_sec")), as.numeric))

输出结果验证

两种方法最终得到的拆分结果一致,所有混乱分隔符都被正确过滤:

latlonglat_deglat_minlat_seclong_deglong_minlong_sec
22ª29'56.06"105º21'37.27"222956.061052137.27
22°29`53.14" | 105°21'29.48" | 22 | 29 | 53.14 | 105 | 21 | 29.48 |
22º30'00.43"105ª21'37.46''22300.431052137.46
105'29'27.17"105°21'39.68"1052927.171052139.68

你拿到度、分、秒三个数值字段后,就可以按照十进制度 = 度 + 分/60 + 秒/3600的公式直接转换为十进制坐标。

内容的提问来源于stack exchange,提问作者Arturo Sbr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 01:09:03