寻求更简洁优雅的R代码实现论文嵌套式前置元数据展示
优化嵌套列表遍历:替代双层循环的R实现
我正在用Obsidian和R清理数十篇论文的YAML属性,已经用上了dplyr、tidyr、purrr、ymlthis这些包,但处理嵌套列表还是有点棘手。下面是一份论文元数据的嵌套列表(转自YAML):
li <- list( aliases = "Al-Kadem2016", type = "paper", title = "Real-Time Estimation of Flow Rate in Dry Gas Wells - A New Field", citationKey = "Al-Kadem2016", id = "SPE-182754-MS", DOI = "https://doi.org/10.2118/182754-MS", ISBN = "", volume = "", url = "", authors = list( "Mohammad S. Al-Kadem", "Mohammad S. Al Dabbous", "Ali S. Al Mashhad", "Hassan A. Alsadah", "Dhafer Al Shehri" ), institutions = list( '[[Saudi Aramco]]', '[[KFUPM]]' ), published = '2016-04-25', conferenceName = "Saudi Arabia Annual Technical Symposium and Exhibition", location = "Dammam, Saudi Arabia", publisher = "SPE", created = '2023-01-26', cdow = "Thursday", downloaded = '2023-01-26', summary = "An empirical correlation was developed to calculate the real-time flow rate in dry gas wells at the surface utilizing the most appropriate parameters: upstream flowing wellhead pressure, downstream flowing wellhead pressure, upstream flowing wellhead temperature, and choke size.", critique = "Shared with VFM team. Contains **Panhandle choke correlation original and modified**. Covers some empirical choke correlations, applicability and limitations.", rating = 4, pages = 10, references = 5, review = "", tags = "", has_glossary = FALSE, cited_by = "", up = '[[Choke Correlation]]', down = "", related = list( '[[Gas Wells]]', '[[Atlas/Concepts/Empirical Correlations]]', '[[Atlas/Concepts/Flow Rate Estimation]]', '[[Atlas/Objects/Venturi Flowmeter]]', '[[Atlas/Objects/Plant Information]]' ), research = "CHK", notes = "what-makes(iiii) + doi + authors + tags(iiii) + snips(iiii) + view + defs(iiiii)", cover = '[[poster-20231026193936.png]]', poster = '[[poster-20231026193936.png]]', pdf_annotated = FALSE, zotero_up = FALSE, zotero_notes = FALSE, zotero_highlights = FALSE, uuid = '20230126065407', added = '2023-01-01', pdf_attached = "spe-000000.pdf", pdf_highlights = FALSE )
我自己写了一段双层循环的代码,能输出符合需求的键值对结果,但代码不够简洁高效,想找更优雅的实现方案,核心是学习嵌套列表的遍历技术,不用ymlthis::as_yml进行格式化。原代码如下:
# get values from frontmatter fm_pluck <- li for(i in 1:length(fm_pluck)) { element <- fm_pluck[i] key <- names(element) values <- unlist(element) n_values <- length(values) cat(sprintf("%2d %2d %-15s", i, n_values, key)) if (n_values > 1) { cat("\n") for (v in 1:n_values) { cat(sprintf("%20d %-12s \n", v, values[v])) } } else { cat(sprintf("%-20s \n", values)) } }
优化方案:用purrr实现函数式遍历
利用你已经在使用的purrr包,可以用函数式编程的方式替代显式循环,代码更简洁易读。
方法1:iwalk+匿名函数
iwalk是purrr中专门用于遍历列表并执行副作用(比如打印输出)的函数,会自动传入元素值、键和索引:
library(purrr) iwalk(li, function(val, key, idx) { vals <- unlist(val) n_vals <- length(vals) # 打印键的头部信息 cat(sprintf("%2d %2d %-15s", idx, n_vals, key)) if (n_vals > 1) { cat("\n") # 遍历多值项,用walk2同时处理索引和值 walk2(seq_len(n_vals), vals, ~cat(sprintf("%20d %-12s \n", .x, .y))) } else { cat(sprintf("%-20s \n", vals)) } })
方法2:紧凑版匿名函数
如果不需要复用处理逻辑,可以把代码写得更紧凑:
library(purrr) iwalk(li, ~{ vals <- unlist(.x) n <- length(vals) cat(sprintf("%2d %2d %-15s", .y, n, names(.x))) if (n > 1) { cat("\n") walk2(seq_len(n), vals, ~cat(sprintf("%20d %-12s \n", ..1, ..2))) } else { cat(sprintf("%-20s \n", vals)) } })
方法3:拆分独立函数(便于复用)
如果需要在多个地方复用处理逻辑,把打印逻辑拆成独立函数更清晰:
library(purrr) print_key_value <- function(val, key, idx) { vals <- unlist(val) n_vals <- length(vals) cat(sprintf("%2d %2d %-15s", idx, n_vals, key)) if (n_vals > 1) { cat("\n") walk2(seq_len(n_vals), vals, ~cat(sprintf("%20d %-12s \n", .x, .y))) } else { cat(sprintf("%-20s \n", vals)) } } # 调用函数遍历列表 iwalk(li, print_key_value)
这些方案都避免了显式的双层循环,利用purrr的函数式工具实现遍历,代码更符合tidyverse生态的风格,也更易维护和扩展。
内容的提问来源于stack exchange,提问作者f0nzie
相关产品推荐
相关产品推荐

