如何用R通过Reddit API精准提取指定帖子的全部评论?
使用R调用Reddit API提取指定帖子全部评论的优化方案
问题背景
正在学习用R调用Reddit API,目标是提取指定帖子(r/Homebrewing下ID为11dd5r3的帖子)的全部纯评论文本。现有代码已完成认证并获取API响应,但存在两个核心问题:
- 用正则提取
<p>标签内容时会混入系统提示等非评论内容 - 不确定当前代码是否能获取帖子的全部评论(含嵌套子评论、分页评论)
现有代码
认证与API请求代码
library(httr) library(jsonlite) # 设置认证参数 auth <- authenticate("some-key1", "some_key2") # 设置用户代理 user_agent <- "my_app/0.1" # 获取访问令牌 response <- POST("https://www.reddit.com/api/v1/access_token", auth = auth, user_agent = user_agent, body = list(grant_type = "password", username = "abc123", password = "123abc")) # 提取访问令牌 access_token <- content(response)$access_token # 构造帖子API请求URL post_id <- "11dd5r3" url <- paste0("https://oauth.reddit.com/r/Homebrewing/comments/", post_id) # 设置请求头 user_agent_string <- "MyApp/1.0" authorization_header <- paste("Bearer ", access_token, sep = "") # 发送请求 response <- GET(url, add_headers(Authorization = authorization_header, `User-Agent` = user_agent_string)) # 解析响应内容 response_json <- rawToChar(response$content)
尝试提取评论的正则代码
final = response_json[1] matches <- gregexpr("<p>(.*?)</p>", final) matches_text <- regmatches(final, matches)[[1]]
遇到的问题
- 正则提取会混入非评论内容,例如:
[213] "<p>Posts are automatically archived after 6 months.</p>" - 无法获取嵌套子评论,且默认API响应仅返回部分评论,无法覆盖全部内容
优化方案
1. 精准解析JSON,直接提取纯评论文本
Reddit API返回的JSON中,评论的纯文本直接存储在body字段中,无需通过HTML标签提取。直接解析JSON结构可彻底避免混入非评论内容:
# 将原始响应解析为可操作的数据结构 parsed_response <- fromJSON(response_json) # 提取顶级评论的纯文本 top_level_comment_text <- parsed_response$data$children$data$body
2. 处理分页与嵌套评论,获取全部内容
Reddit评论采用嵌套结构+分页机制,需通过递归+循环实现全量提取:
① 获取所有分页的顶级评论
get_all_top_comments <- function(post_url, access_token, user_agent) { all_top_comments <- c() after_token <- NULL while(TRUE) { # 构造带分页参数的请求URL request_url <- if(is.null(after_token)) { post_url } else { paste0(post_url, "?after=", after_token) } # 发送请求并解析 response <- GET(request_url, add_headers(Authorization = paste("Bearer", access_token), `User-Agent` = user_agent)) parsed_data <- fromJSON(rawToChar(response$content)) # 提取当前页评论文本 current_comments <- parsed_data$data$children$data$body all_top_comments <- c(all_top_comments, current_comments) # 更新分页标记,无后续内容则退出循环 after_token <- parsed_data$data$after if(is.null(after_token)) break } return(all_top_comments) } # 调用函数获取全部顶级评论 all_top_comments <- get_all_top_comments(url, access_token, user_agent_string)
② 递归提取所有嵌套子评论
extract_nested_comments <- function(replies_raw) { nested_comments <- c() if(!is.null(replies_raw) && !is.na(replies_raw)) { parsed_replies <- fromJSON(replies_raw) if(!is.null(parsed_replies$data$children)) { child_comments <- parsed_replies$data$children$data # 提取当前层级子评论 nested_comments <- c(nested_comments, child_comments$body) # 递归提取更深层级的子评论 for(reply in child_comments$replies) { nested_comments <- c(nested_comments, extract_nested_comments(reply)) } } } return(nested_comments) } # 先获取带replies字段的完整顶级评论数据 full_top_level_data <- parsed_response$data$children$data # 提取所有嵌套子评论 all_nested_comments <- c() for(reply in full_top_level_data$replies) { all_nested_comments <- c(all_nested_comments, extract_nested_comments(reply)) }
③ 合并所有评论
# 合并顶级评论和嵌套子评论 full_comment_list <- c(all_top_comments, all_nested_comments)
3. 可选:清理转义字符
如果评论中存在HTML转义字符(如 ),可使用htmltools包解码:
library(htmltools) clean_comments <- sapply(full_comment_list, html_text)
内容的提问来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

