old.reddit网页抓取项目:如何正确获取帖子描述?
解决Old.Reddit帖子描述抓取失败的问题
问题根源
- 列表页的
div.md是子版块的介绍内容,不是帖子正文,你在fetch_reddit_data里抓取的post_content完全没用,还会造成混淆。 - 帖子页直接用
div.md可能匹配到页面其他区域的内容(比如侧边栏或评论区),需要精准定位到原帖作者的正文区域。
修复方案
1. 修正帖子内容抓取函数(fetch_post_content)
Old.Reddit里,原帖的正文内容嵌套在div.usertext-body下的div.md里,用这个层级定位就能避免误抓其他内容,同时处理帖子无正文的情况(比如纯链接帖):
def fetch_post_content(post_url): response = requests.get(post_url, headers={"User-Agent": "Mozilla/5.0"}) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') # 精准定位原帖作者的正文区域 usertext_body = soup.find("div", class_="usertext-body") if usertext_body: post_content = usertext_body.find("div", class_="md").text.strip() # 如果正文为空,返回提示文本 return post_content if post_content else "该帖子无正文描述" else: return "该帖子无正文描述" else: print(f"Error: Unable to fetch post content. Status code: {response.status_code}") return ""
2. 删除列表页无关代码
在fetch_reddit_data函数里,你抓取了post_content = soup.find_all("div", class_="md")但完全没用到,而且这行还会拿到子版块介绍,直接删除这行即可。
3. 修复文件夹删除函数(可选)
你的reset_folder函数用os.rmdir只能删除空文件夹,非空文件夹会报错,改用shutil.rmtree更合理:
import shutil # 需要在代码顶部导入shutil模块 def reset_folder(base_folder): try: shutil.rmtree(base_folder) print(f"Previous data removed.") except OSError as e: print(f"Error: {e}")
测试验证
修改后,程序会进入对应帖子页面精准抓取原帖正文内容,不会再误拿到子版块介绍,同时也能妥善处理无正文的帖子。
内容的提问来源于stack exchange,提问作者literal-gargoyle
相关产品推荐
相关产品推荐

