You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中对DataFrame每行并行应用osrmRoute函数?

使用osrm包并行计算多对起终点的出行时间与路线

问题说明

现有每行对应一对起终点的数据集,包含字段ID_o、ID_d、longitude_o、latitude_o、longitude_d、latitude_d,需要用osrm包计算出行时间与路线。单条计算已实现,但百万级数据用循环处理效率极低,且遇到NA或请求失败会意外终止,需通过并行方式处理。

已实现的单条计算代码

time.route1 <- osrmRoute(src = mydata[1, c('longitude_o', 'latitude_o')],
                         dst = mydata[1, c('longitude_d', 'latitude_d')],
                         returnclass = "sf")

失败的apply尝试及原因

尝试用apply实现批量处理时报错incorrect number of dimensions,代码如下:

biroute <- function(geofile, ix=1) {
  osrmRoute(src = geofile[ix, c('longitude_o', 'latitude_o')],
            dst = geofile[ix, c('longitude_d', 'latitude_d')])
}
route <- apply(mydata, 1, biroute)

报错原因:apply会将数据框转换为矩阵,导致列名索引失效;且apply的逻辑是将每行数据传递给函数的第一个参数,而当前函数的参数设计(需要整个数据框+索引)不符合apply的调用逻辑。


可行的并行解决方案

方案1:用furrr实现并行(purrr的并行扩展)

先定义单行处理函数,加入NA和错误捕获,避免单个请求失败终止整个任务:

library(purrr)
library(furrr)
library(osrm)
library(dplyr)
library(tibble)

get_route <- function(row) {
  # 检查坐标是否存在NA
  if (any(is.na(c(row$longitude_o, row$latitude_o, row$longitude_d, row$latitude_d)))) {
    return(tibble(ID_o = row$ID_o, ID_d = row$ID_d, duration = NA, distance = NA, geometry = NA))
  }
  # 捕获osrm请求错误
  safe_route <- possibly(osrmRoute, otherwise = NULL)
  route <- safe_route(
    src = c(row$longitude_o, row$latitude_o),
    dst = c(row$longitude_d, row$latitude_d),
    returnclass = "sf"
  )
  # 处理请求失败的情况
  if (is.null(route)) {
    return(tibble(ID_o = row$ID_o, ID_d = row$ID_d, duration = NA, distance = NA, geometry = NA))
  }
  # 合并ID与路线结果
  bind_cols(tibble(ID_o = row$ID_o, ID_d = row$ID_d), select(route, duration, distance, geometry))
}

设置并行会话并执行:

# 根据CPU核心数设置并行线程数,建议不超过核心数的80%
plan(multisession, workers = 4)
# 将数据拆分为单行数据框的列表,并行处理
result_list <- future_map(split(mydata, seq(nrow(mydata))), get_route)
# 合并为最终数据框
result_df <- bind_rows(result_list)
# 关闭并行会话
plan(sequential)

方案2:用foreach+doParallel实现并行

library(foreach)
library(doParallel)
library(osrm)
library(dplyr)
library(tibble)

# 注册并行集群
cl <- makeCluster(4) # 调整为合适的线程数
registerDoParallel(cl)

# 并行循环处理,自动合并结果
result_df <- foreach(i = 1:nrow(mydata), 
                     .combine = bind_rows, 
                     .packages = c("osrm", "dplyr", "tibble")) %dopar% {
  row <- mydata[i, ]
  # 检查NA
  if (any(is.na(c(row$longitude_o, row$latitude_o, row$longitude_d, row$latitude_d)))) {
    return(tibble(ID_o = row$ID_o, ID_d = row$ID_d, duration = NA, distance = NA, geometry = NA))
  }
  # 捕获请求错误
  tryCatch({
    route <- osrmRoute(
      src = c(row$longitude_o, row$latitude_o),
      dst = c(row$longitude_d, row$latitude_d),
      returnclass = "sf"
    )
    bind_cols(tibble(ID_o = row$ID_o, ID_d = row$ID_d), select(route, duration, distance, geometry))
  }, error = function(e) {
    tibble(ID_o = row$ID_o, ID_d = row$ID_d, duration = NA, distance = NA, geometry = NA)
  })
}

# 关闭并行集群
stopCluster(cl)

关键注意事项

  • 若使用公共OSRM服务,并行请求频率过高可能触发限流或封禁,百万级数据建议搭建私有OSRM服务。
  • 必须保留NA/错误捕获逻辑,确保单个请求失败不会导致整个并行任务中断。
  • 并行线程数需根据CPU核心数和OSRM服务承载能力调整,避免资源耗尽。

内容的提问来源于stack exchange,提问作者Alice-W

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 04:20:34