R语言中矩阵转Dataframe失败问题求助
问题:转换API返回数据为DataFrame时异常
我编写了一个调用OpenTapioca API并解析标注的id、label、description和score的函数,但将数据转换为DataFrame时出现异常。目前data_matrix矩阵显示正常,但转换为DataFrame后结果不对。
原函数代码
get_wikidata_links <- function(input_text, minimum_score) { # # Function which takes a character vector of length 1 as input (i.e. all text # needs to be combined into a single character) as well as a minimum certainty # score, and returns a tibble with key information and links to Wikidata # # Input # - input_text: Text input (character) # - minimum_score: Minimum score that every returned entity needs to have # (numeric) # # Output # - top_wikidata_links: Table with the first four columns being 'id', 'label', # 'description', 'score' (tibble) # base_url <- "https://opentapioca.org/api/annotate" r <- GET(base_url, query = list(query = input_text)) data = content(r)$annotations framed = list() vec = list() dummy = 0 for (i in 1:length(data)) { data1 = data[[i]]$tags for (j in 1:length(data1)) { data2 = data1[[j]] if (data2$score>minimum_score) { vec[1] <- data2$id vec[2] <- data2$label vec[3] <- data2$desc vec[4] <- data2$score dummy <- dummy + 1 framed[[dummy]] <- vec } } } data_matrix <- do.call("rbind", framed) top_wikidata_links <- as.data.frame(data_matrix, stringsAsFactors = FALSE) colnames(top_wikidata_links) <- c("ID", "Label", "Description", "Score") return(top_wikidata_links) }
测试代码
# Test 1 text_example_1 <- c("Karl Popper worked at the LSE.") get_wikidata_links(input_text_1, -0.5) # 注:原代码参数名错误,应为text_example_1 # # Hint: The output should be a tibble similar to the one outlined below # # | id | label | description | score | # | "Q81244" | "Karl Popper" | "Austrian-British philosopher of science" | 2.4568285 | # | "Q174570" | "London School of Economics and Political Science" | "university in Westminster, UK" | "1.4685043" | # | "Q171240" | "London Stock Exchange" | "stock exchange in the City of London" | "-0.4124461" | # Test 2 text_example_2 <- c("Claude Shannon studied at the University of Michigan and at MIT.") get_wikidata_links(text_example_2, 0)
问题根源
原代码中vec列表是在循环外部定义的,每次循环只是修改这个列表的元素,而非创建新的列表对象。这导致framed中的所有元素都指向同一个vec,最终所有行都会是最后一次循环的结果,引发DataFrame异常。此外,手动转矩阵再转DataFrame的方式容易丢失数据类型(比如Score会被转为字符型),且效率较低。
修复方案
方案1:修复原循环逻辑
每次符合条件时创建新的列表,避免引用同一对象,同时用bind_rows直接生成tibble:
library(httr) library(tibble) get_wikidata_links <- function(input_text, minimum_score) { base_url <- "https://opentapioca.org/api/annotate" r <- GET(base_url, query = list(query = input_text)) data <- content(r)$annotations framed <- list() dummy <- 0 for (i in 1:length(data)) { data1 <- data[[i]]$tags for (j in 1:length(data1)) { data2 <- data1[[j]] if (data2$score > minimum_score) { # 每次创建独立的列表,避免共享引用 vec <- list( ID = data2$id, Label = data2$label, Description = data2$desc, Score = data2$score ) dummy <- dummy + 1 framed[[dummy]] <- vec } } } # 处理无符合条件结果的情况 if (length(framed) == 0) { return(tibble(ID = character(), Label = character(), Description = character(), Score = numeric())) } top_wikidata_links <- bind_rows(framed) return(top_wikidata_links) }
方案2:用tidyverse简化代码(推荐)
利用purrr包的扁平化和映射操作,避免嵌套循环,代码更简洁可靠:
library(httr) library(tibble) library(purrr) get_wikidata_links <- function(input_text, minimum_score) { base_url <- "https://opentapioca.org/api/annotate" r <- GET(base_url, query = list(query = input_text)) content(r)$annotations %>% map("tags") %>% # 提取每个annotation下的tags列表 flatten() %>% # 扁平化嵌套列表 keep(~ .x$score > minimum_score) %>% # 过滤分数达标项 map_dfr(~ tibble( # 映射并合并为tibble ID = .x$id, Label = .x$label, Description = .x$desc, Score = .x$score )) }
修正后的测试代码
# Test 1 text_example_1 <- c("Karl Popper worked at the LSE.") get_wikidata_links(text_example_1, -0.5) # Test 2 text_example_2 <- c("Claude Shannon studied at the University of Michigan and at MIT.") get_wikidata_links(text_example_2, 0)
内容的提问来源于stack exchange,提问作者Jay
相关产品推荐
相关产品推荐

