使用R读取并合并多个YAML文件时melt函数报错求助
问题描述
尝试用for循环读取同一目录下多个YAML文件并合并成DataFrame,代码如下:
library(yaml) library(reshape2) test_path <- paste0(getwd(), "/tests/") test_files <- list.files(path=test_path, pattern=".yml", all.files=FALSE, full.names=FALSE) df <- data.frame() for (test_file in test_files) out <- yaml.load_file(paste0(test_path, test_file)) x <- melt(out) df <- x
执行x <- melt(out)时触发错误:
"Error in names(object) <- nm : 'names' attribute [1] must be the same length as the vector [0]"
out的结构示例:
> out $`Test-Scenario-1` $`Test-Scenario-1`[[1]] $`Test-Scenario-1`[[1]]$name [1] "Example Test 1 hello" $`Test-Scenario-1`[[1]]$description [1] "low level load test" $`Test-Scenario-1`[[1]]$url [1] "https://www.github.com/" $`Test-Scenario-1`[[1]]$typle [1] "load" $`Test-Scenario-1`[[1]]$thread [1] 7 $`Test-Scenario-1`[[1]]$loop [1] 10 $`Test-Scenario-1`[[1]]$head NULL $`Test-Scenario-1`[[1]]$delay [1] 100
两个示例YAML文件:
hello_test.yml:
Test-Scenario-1: - name: hello Example Test 1 description: hello level load test url: https://www.hello.com/ typle: load thread: 10 loop: 20 head: NO_HEAD delay: 100 Test-Scenario-2: - name: hello Example Test 2 description: hello level load test url: https://www.hello.com/ typle: load thread: 10 loop: 20 head: NO_HEAD delay: 100
example_test.yml:
Test-Scenario-1: - name: Example Test 1 description: low level load test url: https://www.github.com/ typle: load thread: 7 loop: 10 head: NO_HEAD delay: 100 Test-Scenario-2: - name: Example Test 2 description: low level load test url: https://www.github.com/ typle: load thread: 7 loop: 10 head: NO_HEAD delay: 100
错误原因
- for循环语法缺失:你的for循环没有用
{}包裹代码块,导致只有out <- yaml.load_file(...)属于循环体,后续的melt和赋值操作只会在循环结束后执行一次,若最后一个文件的解析结构不符合melt要求,就会触发错误。 - 嵌套列表与
melt的兼容性问题:从out的结构来看,它是多层嵌套列表(顶层是场景名,每个场景下是包含单个配置项的列表,配置项里又有子字段,还存在NULL值)。reshape2::melt默认无法正确处理这种多层嵌套且含空值的结构,会因元素长度不匹配抛出"names属性长度不一致"的错误。 - 数据合并逻辑错误:当前代码中
df <- x是直接覆盖原有数据,而非合并多个文件的结果,完全达不到"合并为DataFrame"的目的。
修正后的代码
改用purrr和dplyr处理更简洁,同时解决上述问题:
library(yaml) library(dplyr) library(purrr) test_path <- file.path(getwd(), "tests") # 用\\.yml$精准匹配后缀,避免匹配到类似.test.yml的文件 test_files <- list.files(path = test_path, pattern = "\\.yml$", full.names = TRUE) # 定义单个YAML文件的处理函数 process_yml <- function(file_path) { out <- yaml.load_file(file_path) # 遍历每个测试场景,转成DataFrame并补充场景名和来源文件信息 map_dfr(out, function(scenario) { # 把NULL值替换为NA,避免后续转DataFrame出错 scenario[[1]] <- map(scenario[[1]], ~ if(is.null(.x)) NA else .x) as.data.frame(scenario[[1]]) %>% mutate(scenario_name = names(out)[which(out == scenario)]) }) %>% mutate(source_file = basename(file_path)) } # 批量处理所有文件并合并成最终DataFrame df <- map_dfr(test_files, process_yml)
内容的提问来源于stack exchange,提问作者HighHill
相关产品推荐
相关产品推荐

