如何将Reddit JSON列表转换为指定格式的Data Frame?
问题
我用以下R代码从Reddit页面提取JSON数据:
library(jsonlite) results <- fromJSON("https://www.reddit.com/r/gardening/comments/1196opl/tree_surgeon_butchered_my_tree_will_it_be_ok/.json") final = results$data
查看输出时,发现它是list格式,但内部有类似表格的数据结构。我尝试用dataframe_list = as.data.frame(final)转成Data Frame,但没得到预期的表格格式。我希望最终得到包含comment_id和comment_text的表格,格式如下:
comment_id comment_text 1 1 I like gardening! 2 2 I dont like to garden! 3 3 its too cold outside? 4 4 try planting something different? 5 5 garden is fun!
注:目标评论文本位于JSON的"body:"与"edited:"标签之间,有没有更优的实现方式?
解决方案
Reddit返回的JSON结构分为两部分:第一部分是原帖数据,第二部分是评论区数据。要提取评论的ID和内容,需要定位到评论所在的层级:
定位评论数据层级
评论数据存储在results[[2]]$data$children中,每个子元素对应单条评论,核心字段都在data子节点下。提取目标字段并构建Data Frame
直接从评论节点中提取id(作为comment_id)和body(作为comment_text),组合成表格:
library(jsonlite) # 读取Reddit JSON数据 results <- fromJSON("https://www.reddit.com/r/gardening/comments/1196opl/tree_surgeon_butchered_my_tree_will_it_be_ok/.json") # 提取评论区的原始数据 comments_data <- results[[2]]$data$children # 构建目标表格 comment_table <- data.frame( comment_id = sapply(comments_data, function(x) x$data$id), comment_text = sapply(comments_data, function(x) x$data$body), stringsAsFactors = FALSE ) # 查看结果 head(comment_table)
- 更简洁的实现(借助purrr包)
如果安装了purrr包,可以用更简洁的语法提取字段:
library(purrr) library(tibble) comment_table <- map_dfr(comments_data, ~{ tibble( comment_id = .x$data$id, comment_text = .x$data$body ) })
关键说明
- 直接用
as.data.frame(final)无效的原因:results$data对应的是原帖的元数据,并非评论区的结构化数据,评论区在JSON的第二个顶级元素中。 - 字段对应关系:JSON中的
id就是评论唯一标识,body就是评论内容,无需依赖edited:字段,直接提取body即可得到目标文本。
内容的提问来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

