You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将捆绑经纬度的字符串向量转换为R语言数据框?

提取位置字符串中的经纬度变量

需要从包含lat:和lng:格式的字符串中,拆分出纬度(latitude)和经度(longitude)作为单独变量,原始数据如下:

ID <- c(1, 2, 3)
location_1 <- c("lat:10.1234567,lng:-70.1234567", "lat:20.1234567891234,lng:-80.1234567891234", "lat:30.1234567,lng:-90.1234567")

df <- data.frame(ID, location_1)

# 数据预览
#   ID                          location_1
# 1  1        lat:10.1234567,lng:-70.1234567
# 2  2 lat:20.1234567891234,lng:-80.1234567891234
# 3  3        lat:30.1234567,lng:-90.1234567

期望得到的结果:

#   ID       latitude        longitude
# 1  1      10.1234567      -70.1234567
# 2  2 20.1234567891234 -80.1234567891234
# 3  3      30.1234567      -90.1234567

之前方法的问题

  • 方法一:正则未匹配负号,且[[1]]仅提取第一行数据,同时正则中的,\\s和(?=\\))与原始数据格式不匹配(无空格、无括号),导致返回NA。
  • 方法二:正则匹配的是括号包裹的数值,但原始数据无括号结构,完全不匹配。

正确解决方法

方法1:使用tidyr包(最简便)

利用separate_wider_regex直接匹配lat:和lng:后的数值:

library(tidyr)

df_result <- df %>%
  separate_wider_regex(
    location_1,
    patterns = list(
      latitude = "(?<=lat:)[-0-9.]+",
      ignore = ",lng:",
      longitude = "[-0-9.]+"
    )
  ) %>%
  mutate(across(c(latitude, longitude), as.numeric))

方法2:使用stringr包

用正则正向预查提取lat:和lng:后的数值:

library(stringr)

df$latitude <- as.numeric(str_extract(df$location_1, "(?<=lat:)[-0-9.]+"))
df$longitude <- as.numeric(str_extract(df$location_1, "(?<=lng:)[-0-9.]+"))

df_result <- df %>% select(ID, latitude, longitude)

方法3:Base R实现

通过strsplit拆分字符串后提取数值:

# 拆分每个location字符串
location_split <- strsplit(df$location_1, ",")

# 提取纬度和经度
df$latitude <- sapply(location_split, function(x) {
  as.numeric(sub("lat:", "", x[1]))
})
df$longitude <- sapply(location_split, function(x) {
  as.numeric(sub("lng:", "", x[2]))
})

df_result <- df[, c("ID", "latitude", "longitude")]

内容的提问来源于stack exchange,提问作者Jennifer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 01:31:03