You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup爬取Reddit用户嵌套评论?现有代码仅能获取页面

如何修改代码爬取Reddit r/CryptoCurrency的用户评论内容

嘿,我看你现在的代码只能拿到r/CryptoCurrency的帖子列表页内容,但还没深入到具体帖子里抓取用户评论对吧?其实只要多两步操作就能搞定:先从列表页提取每个帖子的详情链接,再进入详情页解析评论内容。我给你调整了代码,顺便把关键点说清楚:

核心思路拆解

  • 第一步:从帖子列表页,提取每个帖子的详情页完整链接
  • 第二步:对每个详情页发起请求,解析页面中的评论内容(包括作者、内容、时间等信息)
  • 第三步:处理翻页逻辑,确保能抓取多页帖子的评论

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import time
import json

# 模拟浏览器请求头,避免被Reddit识别为爬虫
headers = {'User-agent': 'My Crypto Comment Crawler/1.0'}
# 用来存储所有评论的列表
all_comments = []
# 翻页用的after参数,初始为空
after_id = ""

# 循环爬取4页(每页25条帖子)
for page in range(4):
    # 构造列表页URL
    list_url = f"https://www.reddit.com/r/CryptoCurrency/?count={page*25}&after={after_id}"
    response = requests.get(list_url, headers=headers)
    soup = BeautifulSoup(response.content, 'html.parser')
    
    # 提取当前页所有帖子的详情链接
    post_title_links = soup.select('a[data-testid="post-title"]')
    for link in post_title_links:
        post_full_url = f"https://www.reddit.com{link['href']}"
        print(f"正在处理帖子: {post_full_url}")
        
        # 请求帖子详情页,加个延迟避免触发反爬
        time.sleep(2)
        post_response = requests.get(post_full_url, headers=headers)
        post_soup = BeautifulSoup(post_response.content, 'html.parser')
        
        # 提取当前帖子的所有评论(包括嵌套回复)
        comment_containers = post_soup.select('div[data-testid="comment"]')
        for comment in comment_containers:
            # 提取评论作者,处理已删除账号的情况
            author_elem = comment.select_one('a[data-testid="comment-author"]')
            author = author_elem.text if author_elem else "[已删除]"
            
            # 提取评论内容
            content_elem = comment.select_one('div[data-testid="comment-content"]')
            content = content_elem.get_text(strip=True) if content_elem else "[无内容]"
            
            # 提取评论发布时间
            time_elem = comment.select_one('time')
            post_time = time_elem['datetime'] if time_elem else "未知时间"
            
            # 将评论信息存入列表
            all_comments.append({
                '帖子链接': post_full_url,
                '评论作者': author,
                '评论内容': content,
                '发布时间': post_time
            })
            print(f"作者: {author} | 内容预览: {content[:60]}...")
    
    # 更新翻页用的after_id,准备爬取下一页
    next_page_link = soup.select_one('span[data-testid="next-button"] a')
    if next_page_link:
        after_id = next_page_link['href'].split('after=')[1]
        print(f"准备爬取下一页,after参数: {after_id}")
    else:
        print("已无更多页面,停止爬取")
        break

# 将所有评论保存为JSON文件,方便后续分析
with open('crypto_currency_comments.json', 'w', encoding='utf-8') as f:
    json.dump(all_comments, f, ensure_ascii=False, indent=4)

print(f"爬取完成,共获取{len(all_comments)}条评论,已保存到crypto_currency_comments.json")

关键细节说明

  • 翻页参数修正:你原来代码里用id_list[-1]是不对的,Reddit的after参数需要帖子的完整标识(比如t3_xxxxxx),所以我们从下一页按钮的链接里提取这个参数。
  • 评论定位:用data-testid属性定位元素更可靠,因为Reddit的CSS类名经常变动,但测试ID一般保持稳定。
  • 反爬处理:加入time.sleep(2)控制请求频率,避免被Reddit临时封禁IP;同时用自定义的User-agent模拟正常浏览器请求。
  • 嵌套评论处理:代码里的select('div[data-testid="comment"]')会抓取所有层级的评论(包括一级评论和回复),如果只想抓一级评论,可以把选择器改成div[data-testid="comment"]:not([data-testid*="reply"])。

内容的提问来源于stack exchange,提问作者Jake Park

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:08:07