如何在purrr::map中跳过无数据eventid并继续处理其余条目?
如何跳过无数据的eventid继续执行函数
我有一个返回eventid列表的函数,列表内容如下:
[1] "314cdd9cecb9b66219c944996f0249b2" "ab26545fc28c693c52329db2a68a06a9" "818b7fece8a6f82cecf0b4ef38f903d3" [4] "9427475e5b454a44a4d9cc365d72e416" "3ee7569ac3f38bd5d62ca03b13ee1344" "35c8c545c966fd2ad738dcccd02a6d8f" [7] "f2aa3876acabca48a66ea6eae3d9f1b9" "63308e00869e2009bc6597edbdaf0c99" "96e9ac4414cf1fb0ef4164d710c420c8" [10] "e6b22267452d3dd4d99fa62ff71f7fcc" "8d78600b0a7ab6a7f15c5f935cbace68" "80d27fa88ccd30a960caacce5dfac049" [13] "c95982844161e998ab02c1eab506902d"
随后我使用以下代码,基于上述列表中的eventid执行自定义函数my_func:
purrr::map(event_ids, ~my_func(sport = "basketball_nba", eventId = .x))
正常情况下,该代码会返回类似如下的tibble结果:
my_func('basketball_nba', '314cdd9cecb9b66219c944996f0249b2') # A tibble: 475 × 13 id sport_key sport…¹ comme…² home_…³ away_…⁴ bookm…⁵ title key last_…⁶ name price point <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <int> <dbl> 1 314cdd9cecb9b66219c944996f0249b2 basketbal… NBA 2023-0… Charlo… Miami … draftk… Draf… alte… 2023-0… Char… -475 15.5
但当某个eventid无对应数据时,会抛出如下错误:
> nba_alt_lines <- purrr::map(event_ids, ~my_func(sport = "basketball_nba", eventId = .x)) Error in `purrr::map()`: ℹ In index: 7. Caused by error in `rename()`: ! Can't rename columns that don't exist. ✖ Column `key` doesn't exist. Run `rlang::last_error()` to see where the error occurred. > rlang::last_error() <error/purrr_error_indexed> Error in `purrr::map()`: ℹ In index: 7. Caused by error in `rename()`: ! Can't rename columns that don't exist. ✖ Column `key` doesn't exist. --- Backtrace: 1. purrr::map(event_ids, ~my_func(sport = "basketball_nba", eventId = .x)) 12. dplyr:::rename.data.frame(., bookmaker_key = "key")
请问是否可以通过编程方式跳过无数据的eventid,继续处理其余有效的eventid?
解决方案
方法1:使用purrr::safely()包装函数
safely()会为每个函数调用返回包含result和error的列表,出错时result为NULL,error存储错误详情。可以借此过滤出成功的结果:
# 用safely包装my_func,保留错误信息 safe_my_func <- purrr::safely(my_func) # 批量执行所有eventid results <- purrr::map(event_ids, ~safe_my_func(sport = "basketball_nba", eventId = .x)) # 提取成功的结果,过滤掉NULL值 successful_results <- purrr::keep(results, ~!is.null(.x$result)) %>% purrr::map("result") # 合并所有成功结果为单个tibble(可选) combined_tibble <- dplyr::bind_rows(successful_results)
方法2:使用purrr::possibly()返回默认值
possibly()可以指定函数出错时返回的默认值,比如空tibble,后续直接合并即可:
# 定义出错时返回空tibble(建议和正常结果列结构一致,避免合并报错) possibly_my_func <- purrr::possibly( my_func, otherwise = tibble::tibble( id = character(), sport_key = character(), sport_title = character(), commence_time = character(), home_team = character(), away_team = character(), bookmaker_key = character(), bookmaker_title = character(), market_key = character(), market_last_update = character(), outcome_name = character(), price = integer(), point = numeric() ) ) # 批量执行,出错时返回预设空tibble results <- purrr::map(event_ids, ~possibly_my_func(sport = "basketball_nba", eventId = .x)) # 合并结果并过滤空行 combined_tibble <- dplyr::bind_rows(results) %>% dplyr::filter(!dplyr::if_all(dplyr::everything(), is.na))
方法3:手动用tryCatch捕获错误
如果不想依赖purrr的包装函数,可以在匿名函数中直接用tryCatch捕获错误,跳过出错的eventid:
results <- purrr::map(event_ids, function(id) { tryCatch( # 尝试执行函数 my_func(sport = "basketball_nba", eventId = id), # 捕获错误时执行的逻辑 error = function(e) { # 可选:打印错误信息,方便排查 message(sprintf("跳过eventid %s: %s", id, e$message)) # 返回NULL,后续过滤 NULL } ) }) # 过滤NULL结果并合并为单个tibble combined_tibble <- dplyr::bind_rows(purrr::compact(results))
内容的提问来源于stack exchange,提问作者Aaron Morris
相关产品推荐
相关产品推荐

