You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R通过Reddit API精准提取指定帖子的全部评论?

使用R调用Reddit API提取指定帖子全部评论的优化方案

问题背景

正在学习用R调用Reddit API,目标是提取指定帖子(r/Homebrewing下ID为11dd5r3的帖子)的全部纯评论文本。现有代码已完成认证并获取API响应,但存在两个核心问题:

  1. 用正则提取<p>标签内容时会混入系统提示等非评论内容
  2. 不确定当前代码是否能获取帖子的全部评论(含嵌套子评论、分页评论)

现有代码

认证与API请求代码

library(httr)
library(jsonlite)

# 设置认证参数
auth <- authenticate("some-key1", "some_key2")

# 设置用户代理
user_agent <- "my_app/0.1"

# 获取访问令牌
response <- POST("https://www.reddit.com/api/v1/access_token",
                 auth = auth,
                 user_agent = user_agent,
                 body = list(grant_type = "password",
                             username = "abc123",
                             password = "123abc"))

# 提取访问令牌
access_token <- content(response)$access_token

# 构造帖子API请求URL
post_id <- "11dd5r3"
url <- paste0("https://oauth.reddit.com/r/Homebrewing/comments/", post_id)

# 设置请求头
user_agent_string <- "MyApp/1.0"
authorization_header <- paste("Bearer ", access_token, sep = "")

# 发送请求
response <- GET(url, add_headers(Authorization = authorization_header, `User-Agent` = user_agent_string))

# 解析响应内容
response_json <- rawToChar(response$content)

尝试提取评论的正则代码

final = response_json[1]
matches <- gregexpr("<p>(.*?)</p>", final)
matches_text <- regmatches(final, matches)[[1]]

遇到的问题

  • 正则提取会混入非评论内容,例如:
    [213] "<p>Posts&#32;are&#32;automatically&#32;archived&#32;after&#32;6&#32;months.</p>"
    
  • 无法获取嵌套子评论,且默认API响应仅返回部分评论,无法覆盖全部内容

优化方案

1. 精准解析JSON,直接提取纯评论文本

Reddit API返回的JSON中,评论的纯文本直接存储在body字段中,无需通过HTML标签提取。直接解析JSON结构可彻底避免混入非评论内容:

# 将原始响应解析为可操作的数据结构
parsed_response <- fromJSON(response_json)

# 提取顶级评论的纯文本
top_level_comment_text <- parsed_response$data$children$data$body

2. 处理分页与嵌套评论,获取全部内容

Reddit评论采用嵌套结构+分页机制,需通过递归+循环实现全量提取:

① 获取所有分页的顶级评论

get_all_top_comments <- function(post_url, access_token, user_agent) {
  all_top_comments <- c()
  after_token <- NULL
  
  while(TRUE) {
    # 构造带分页参数的请求URL
    request_url <- if(is.null(after_token)) {
      post_url
    } else {
      paste0(post_url, "?after=", after_token)
    }
    
    # 发送请求并解析
    response <- GET(request_url, 
                    add_headers(Authorization = paste("Bearer", access_token), 
                                `User-Agent` = user_agent))
    parsed_data <- fromJSON(rawToChar(response$content))
    
    # 提取当前页评论文本
    current_comments <- parsed_data$data$children$data$body
    all_top_comments <- c(all_top_comments, current_comments)
    
    # 更新分页标记,无后续内容则退出循环
    after_token <- parsed_data$data$after
    if(is.null(after_token)) break
  }
  
  return(all_top_comments)
}

# 调用函数获取全部顶级评论
all_top_comments <- get_all_top_comments(url, access_token, user_agent_string)

② 递归提取所有嵌套子评论

extract_nested_comments <- function(replies_raw) {
  nested_comments <- c()
  
  if(!is.null(replies_raw) && !is.na(replies_raw)) {
    parsed_replies <- fromJSON(replies_raw)
    if(!is.null(parsed_replies$data$children)) {
      child_comments <- parsed_replies$data$children$data
      
      # 提取当前层级子评论
      nested_comments <- c(nested_comments, child_comments$body)
      
      # 递归提取更深层级的子评论
      for(reply in child_comments$replies) {
        nested_comments <- c(nested_comments, extract_nested_comments(reply))
      }
    }
  }
  
  return(nested_comments)
}

# 先获取带replies字段的完整顶级评论数据
full_top_level_data <- parsed_response$data$children$data

# 提取所有嵌套子评论
all_nested_comments <- c()
for(reply in full_top_level_data$replies) {
  all_nested_comments <- c(all_nested_comments, extract_nested_comments(reply))
}

③ 合并所有评论

# 合并顶级评论和嵌套子评论
full_comment_list <- c(all_top_comments, all_nested_comments)

3. 可选:清理转义字符

如果评论中存在HTML转义字符(如&#32;),可使用htmltools包解码:

library(htmltools)
clean_comments <- sapply(full_comment_list, html_text)

内容的提问来源于stack exchange,提问作者stats_noob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 10:52:58