You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup遍历Div提取Reddit帖子评论段落内容?

问题解决及原因分析

一、第二次代码报错的解决方法

你碰到的AttributeError是因为soup.findAll()返回的是元素列表(ResultSet对象),不能直接在列表上调用find()方法,必须先遍历列表里的每一个元素。修正后的代码如下:

# 先获取评论树列表,遍历每个评论树
every_review = soup.findAll('shreddit-comment-tree', {'id':'comment-tree'})
for tree in every_review:
    # 在每个评论树里查找评论内容容器
    comment_containers = tree.find_all('div', {'class':'md pl-[28px] xs:pl-xl text-14'})
    for container in comment_containers:
        # 定位内容div并提取p标签文本
        content_div = container.find('div', {'id':'-post-rtjson-content'})
        if content_div:
            p_tag = content_div.find('p')
            if p_tag:
                print(p_tag.text)

如果要生成你需要的列表格式,可改成:

comment_list = []
every_review = soup.findAll('shreddit-comment-tree', {'id':'comment-tree'})
for tree in every_review:
    comment_containers = tree.find_all('div', {'class':'md pl-[28px] xs:pl-xl text-14'})
    for container in comment_containers:
        content_div = container.find('div', {'id':'-post-rtjson-content'})
        if content_div and content_div.find('p'):
            comment_list.append(content_div.find('p').text)

print(comment_list)  # 输出格式:[sent1, sent2, sent3...]

二、首次代码无输出的原因

  1. 定位元素不匹配:你指定的py-0 xs:ml-xs inline-block max-w-full这个类名,在当前抓取的静态HTML里没有对应的元素,自然遍历不到内容。
  2. 动态加载限制:Reddit的评论大多是通过JavaScript动态渲染的,BeautifulSoup只能解析页面初始加载的静态HTML,抓不到JS后续生成的评论内容。这种情况需要用Selenium、Playwright等工具先渲染页面,再提取HTML。

额外处理方案

如果静态抓取无效,试试用Selenium模拟浏览器加载页面:

from selenium import webdriver
from bs4 import BeautifulSoup
import time

driver = webdriver.Chrome()
driver.get("你的Reddit帖子URL")
time.sleep(3)  # 等待页面加载完成

soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 用修正后的代码提取评论
comment_list = []
every_review = soup.findAll('shreddit-comment-tree', {'id':'comment-tree'})
for tree in every_review:
    comment_containers = tree.find_all('div', {'class':'md pl-[28px] xs:pl-xl text-14'})
    for container in comment_containers:
        content_div = container.find('div', {'id':'-post-rtjson-content'})
        if content_div and content_div.find('p'):
            comment_list.append(content_div.find('p').text)

print(comment_list)

内容的提问来源于stack exchange,提问作者LaurenLai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 05:24:56