You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取论坛帖子生成嵌套字典以获取回复及用户信息?

论坛帖子爬取与嵌套字典实现方案

以下是针对你需求的具体实现步骤,基于你已有的初始代码修改:

1. 准备工作

确保已安装requests和bs4库,未安装的话执行:

pip install requests beautifulsoup4

2. 完善爬取逻辑

修改你的scrape_thread函数,逐个提取目标字段并组装成字典:

import requests
from bs4 import BeautifulSoup

url = "https://www.dentistry-forums.com/threads/extracting-only-one-of-three-tooth-roots-procedure-name.23919/"
response = requests.get(url)
soup = BeautifulSoup(response.content, "html.parser")

def scrape_thread(bowl):
    thread_data = []
    
    for spoon in bowl.select(".message-inner article"):
        # 提取用户名
        user = spoon.select_one(".username").text.strip()
        
        # 提取账号类型,处理无类型的情况
        account_type_elem = spoon.select_one(".userTitle")
        account_type = account_type_elem.text.strip() if account_type_elem else "普通用户"
        
        # 判断是否为原帖发布者
        is_op = bool(spoon.select_one(".message-userDetails .label.label--primary"))
        
        # 提取发布时间
        date_time = spoon.select_one(".u-dt").get("title").strip()
        
        # 提取帖子内容
        response_content = spoon.select_one(".bbWrapper").text.strip()
        
        # 提取帖子编号,修正示例中重复的Response键名
        post_number = spoon.select_one(".message-attribution-opposite .u-concealed").text.strip()
        
        # 提取点赞数,处理无点赞的情况
        votes_elem = spoon.select_one(".voteResult")
        votes = int(votes_elem.text.strip()) if votes_elem else 0
        
        # 提取引用内容,处理无引用的情况
        reference_elem = spoon.select_one(".bbCodeBlock--quote .bbCodeBlock-content")
        reference = reference_elem.text.strip() if reference_elem else ""
        
        # 组装目标字典
        post_dict = {
            'User': user,
            'Reference': reference,
            'Response': response_content,
            'Account Type': account_type,
            'Original Poster': is_op,
            'DateTime': date_time,
            'Votes': votes,
            'Post Number': post_number
        }
        
        thread_data.append(post_dict)
    
    return thread_data

# 调用函数查看结果
result = scrape_thread(soup)
for item in result:
    print(item)
    print("---")

3. 关键细节说明

  • 空值处理:对可能不存在的元素(比如引用、点赞)做判断,避免程序报错。
  • 原帖标识:通过页面中label--primary标签判断是否为原帖发布者。
  • 键名修正:示例中重复的Response键名改为Post Number,避免字典键冲突。
  • 嵌套结构:最终返回的thread_data是列表嵌套字典的结构,每个内层字典对应一条用户的提问或回复。

内容的提问来源于stack exchange,提问作者jack gell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 03:10:38