You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取5ch单帖内容失败,寻求技术解决方案

解决5ch论坛爬虫单帖内容爬取问题

问题根源

你当前代码只提取了线程下的第一个<dd>标签,而5ch的线程结构中,每个帖子由一组<dt>(含ID、日期)和<dd>(内容)组成,一个线程包含多组这样的标签对。直接用find('dd')会获取整个线程的所有内容,而非单帖内容。

修正方案

  1. 遍历线程内所有<dt>标签,逐个匹配对应的<dd>(每个dt的下一个兄弟节点就是对应帖子的dd)
  2. 对每个帖子单独提取ID、日期,避免批量正则导致的信息混乱
  3. 调整数据结构,让每条记录对应单帖的完整信息

修改后的代码

import requests
import sys
import time
import re
from bs4 import BeautifulSoup

board_name = 'campus'
upperframe = []

print(f'processing: {board_name} board from 5Chan')
url = f'https://kizuna.5ch.net/{board_name}/?v=pc'
print(url)

headers = {'user-agent': 'Chrome/104.0.0.0'}

try:
    page = requests.get(url, headers=headers)
    page.encoding = 'shift_jis'
except Exception as e:
    error_type, error_obj, error_info = sys.exc_info()
    print('ERROR FOR LINK:', url)
    print(error_type, 'Line:', error_info.tb_lineno)
else:
    time.sleep(5)
    soup = BeautifulSoup(page.text, 'html.parser')
    threads = soup.find_all('div', attrs={'class': 'THREAD_CONTENTS'})
    print(f'找到{len(threads)}个线程')

    for thread in threads:
        # 获取线程标题
        thread_title = thread.find("h3", attrs={'class': 'thread_title'}).find('span').text.strip()
        # 获取当前线程下的所有帖子头部节点
        post_dts = thread.find("dl", attrs={'class': 'thread'}).find_all('dt')
        
        for dt in post_dts:
            # 提取帖子的ID和日期
            dt_text = dt.text.strip()
            post_id = re.search(r'ID:[A-Za-z0-9/+]+', dt_text).group() if re.search(r'ID:[A-Za-z0-9/+]+', dt_text) else None
            post_date = re.search(r'\d{4}/\d{2}/\d{2}', dt_text).group() if re.search(r'\d{4}/\d{2}/\d{2}', dt_text) else None
            # 获取对应帖子的内容
            post_content = dt.find_next_sibling('dd').text.strip()
            # 将单帖信息存入列表
            upperframe.append({
                'thread_title': thread_title,
                'post_id': post_id,
                'post_date': post_date,
                'post_content': post_content
            })

# 打印采集结果数量
print(f'共采集{len(upperframe)}条帖子')

说明

  • 改用find_all('dt')获取线程内所有帖子的头部节点,再通过find_next_sibling('dd')精准匹配对应帖子的内容
  • 每个帖子单独处理ID和日期,避免批量拼接后正则导致的信息错位
  • 用字典存储单帖信息,结构更清晰,后续处理更方便

内容的提问来源于stack exchange,提问作者MichaelWCC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 16:55:15