You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

old.reddit网页抓取项目:如何正确获取帖子描述?

解决Old.Reddit帖子描述抓取失败的问题

问题根源

  1. 列表页的div.md是子版块的介绍内容,不是帖子正文,你在fetch_reddit_data里抓取的post_content完全没用,还会造成混淆。
  2. 帖子页直接用div.md可能匹配到页面其他区域的内容(比如侧边栏或评论区),需要精准定位到原帖作者的正文区域。

修复方案

1. 修正帖子内容抓取函数(fetch_post_content)

Old.Reddit里,原帖的正文内容嵌套在div.usertext-body下的div.md里,用这个层级定位就能避免误抓其他内容,同时处理帖子无正文的情况(比如纯链接帖):

def fetch_post_content(post_url):
    response = requests.get(post_url, headers={"User-Agent": "Mozilla/5.0"})
    if response.status_code == 200:
        soup = BeautifulSoup(response.text, 'html.parser')
        # 精准定位原帖作者的正文区域
        usertext_body = soup.find("div", class_="usertext-body")
        if usertext_body:
            post_content = usertext_body.find("div", class_="md").text.strip()
            # 如果正文为空,返回提示文本
            return post_content if post_content else "该帖子无正文描述"
        else:
            return "该帖子无正文描述"
    else:
        print(f"Error: Unable to fetch post content. Status code: {response.status_code}")
        return ""

2. 删除列表页无关代码

在fetch_reddit_data函数里,你抓取了post_content = soup.find_all("div", class_="md")但完全没用到,而且这行还会拿到子版块介绍,直接删除这行即可。

3. 修复文件夹删除函数(可选)

你的reset_folder函数用os.rmdir只能删除空文件夹,非空文件夹会报错,改用shutil.rmtree更合理:

import shutil  # 需要在代码顶部导入shutil模块

def reset_folder(base_folder):
    try:
        shutil.rmtree(base_folder)
        print(f"Previous data removed.")
    except OSError as e:
        print(f"Error: {e}")

测试验证

修改后,程序会进入对应帖子页面精准抓取原帖正文内容,不会再误拿到子版块介绍,同时也能妥善处理无正文的帖子。

内容的提问来源于stack exchange,提问作者literal-gargoyle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 17:06:02