使用BeautifulSoup爬取5ch单帖内容失败,寻求技术解决方案
解决5ch论坛爬虫单帖内容爬取问题
问题根源
你当前代码只提取了线程下的第一个<dd>标签,而5ch的线程结构中,每个帖子由一组<dt>(含ID、日期)和<dd>(内容)组成,一个线程包含多组这样的标签对。直接用find('dd')会获取整个线程的所有内容,而非单帖内容。
修正方案
- 遍历线程内所有
<dt>标签,逐个匹配对应的<dd>(每个dt的下一个兄弟节点就是对应帖子的dd) - 对每个帖子单独提取ID、日期,避免批量正则导致的信息混乱
- 调整数据结构,让每条记录对应单帖的完整信息
修改后的代码
import requests import sys import time import re from bs4 import BeautifulSoup board_name = 'campus' upperframe = [] print(f'processing: {board_name} board from 5Chan') url = f'https://kizuna.5ch.net/{board_name}/?v=pc' print(url) headers = {'user-agent': 'Chrome/104.0.0.0'} try: page = requests.get(url, headers=headers) page.encoding = 'shift_jis' except Exception as e: error_type, error_obj, error_info = sys.exc_info() print('ERROR FOR LINK:', url) print(error_type, 'Line:', error_info.tb_lineno) else: time.sleep(5) soup = BeautifulSoup(page.text, 'html.parser') threads = soup.find_all('div', attrs={'class': 'THREAD_CONTENTS'}) print(f'找到{len(threads)}个线程') for thread in threads: # 获取线程标题 thread_title = thread.find("h3", attrs={'class': 'thread_title'}).find('span').text.strip() # 获取当前线程下的所有帖子头部节点 post_dts = thread.find("dl", attrs={'class': 'thread'}).find_all('dt') for dt in post_dts: # 提取帖子的ID和日期 dt_text = dt.text.strip() post_id = re.search(r'ID:[A-Za-z0-9/+]+', dt_text).group() if re.search(r'ID:[A-Za-z0-9/+]+', dt_text) else None post_date = re.search(r'\d{4}/\d{2}/\d{2}', dt_text).group() if re.search(r'\d{4}/\d{2}/\d{2}', dt_text) else None # 获取对应帖子的内容 post_content = dt.find_next_sibling('dd').text.strip() # 将单帖信息存入列表 upperframe.append({ 'thread_title': thread_title, 'post_id': post_id, 'post_date': post_date, 'post_content': post_content }) # 打印采集结果数量 print(f'共采集{len(upperframe)}条帖子')
说明
- 改用
find_all('dt')获取线程内所有帖子的头部节点,再通过find_next_sibling('dd')精准匹配对应帖子的内容 - 每个帖子单独处理ID和日期,避免批量拼接后正则导致的信息错位
- 用字典存储单帖信息,结构更清晰,后续处理更方便
内容的提问来源于stack exchange,提问作者MichaelWCC
相关产品推荐
相关产品推荐

