You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup爬取分页页面:解决URL直接访问404问题

解决分页爬取中直接访问URL返回404的问题

这种场景本质是网站反爬机制在起作用,常见原因包括验证请求上下文、会话状态,或者限制直接跳页的行为。以下是针对性的解决方法和代码调整:

核心问题分析

  • 网站检查Referer请求头:直接输入URL时Referer为空或不符合规则,服务器判定为非正常访问返回404
  • 会话Cookie验证:网站需要保持连续会话,单独请求高页码时缺少必要的Cookie
  • 分页参数并非简单递增:手动构造的pageNumber参数可能未包含网站要求的动态验证信息(如签名、token)
  • 反爬限制:网站禁止直接跳转到高页码,必须按顺序从页面中获取下一页链接

具体解决方案及代码修改

1. 使用Session维持会话

改用requests.Session()管理请求,自动维持Cookie,模拟浏览器的连续会话状态。

2. 从页面中解析下一页链接,而非手动构造URL

放弃手动拼接pageNumber的方式,通过当前页面的分页按钮提取真实的下一页URL,确保参数符合网站要求。

3. 添加Referer请求头

在每次请求时设置Referer为上一页的URL,模拟正常的页面跳转行为。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import csv
import time
import logging
from urllib.parse import urljoin

# 配置日志
logging.basicConfig(level=logging.INFO,
                    format='%(asctime)s - %(levelname)s - %(message)s')

# 初始化Session,维持会话状态
session = requests.Session()
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:89.0) Gecko/20100101 Firefox/89.0"
}
session.headers.update(headers)


def fetch_pagination_url(current_url):
    response = session.get(current_url)
    if response.status_code == 200:
        soup = BeautifulSoup(response.text, 'html.parser')
        next_page = soup.select_one('.pagination__link--next')
        
        if next_page and next_page.get('href'):
            next_url = urljoin("https://mywebsite.com/", next_page.get('href'))
            logging.info(f"获取下一页URL: {next_url}")
            return next_url, response.text  # 返回下一页URL和当前页HTML用于解析数据
    logging.error(f"获取下一页链接失败,当前URL: {current_url}")
    return None, None


def parse_data(html):
    soup = BeautifulSoup(html, 'html.parser')
    courses = soup.find_all('div', class_='grid__item')
    data = []

    for course in courses:
        course_data = {}

        title_element = course.find('h2', class_=['h4', 'course-title'])
        course_data['title'] = title_element.text.strip() if title_element else None

        provider_element = course.find('div', class_='provider')
        course_data['provider'] = provider_element.text.strip() if provider_element else None

        location_element = course.find('div', class_='location')
        course_data['location'] = location_element.text.strip() if location_element else None

        data.append(course_data)

    return data


def save_data(data, filename):
    with open(filename, 'a', newline='', encoding='utf-8') as csvfile:
        fieldnames = ['title', 'provider', 'location']
        writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
        for row in data:
            writer.writerow(row)


def main():
    page = 1
    output_filename = 'output.csv'
    # 初始化起始URL(从第一页开始)
    current_url = "https://website.com/coursedisplay/results/courses?studyYear=2023&destination=Undergraduate&postcodeDistanceSystem=imperial&pageNumber=1&sort=MostRelevant&clearingPreference=None"

    # 写入CSV表头
    with open(output_filename, 'w', newline='', encoding='utf-8') as csvfile:
        fieldnames = ['title', 'provider', 'location']
        writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
        writer.writeheader()

    while current_url:
        # 获取下一页URL和当前页HTML
        next_url, html = fetch_pagination_url(current_url)
        if html:
            data = parse_data(html)
            save_data(data, output_filename)
            logging.info(f"已爬取第 {page} 页")
            page += 1
        else:
            logging.info("没有更多页面或请求失败,停止爬取")
            break
        
        # 更新当前URL为下一页地址,并设置Referer头
        current_url = next_url
        if current_url:
            session.headers['Referer'] = current_url  # 添加Referer头
        
        # 延迟避免服务器压力
        time.sleep(10)


if __name__ == '__main__':
    main()

额外注意事项

  • 不要跳过页码请求,必须按顺序从第一页开始依次爬取,避免触发网站的反爬限制
  • 若网站使用JavaScript渲染分页,可能需要改用selenium或playwright模拟浏览器行为,直接执行点击分页按钮的操作
  • 可适当调整延迟时间,或添加随机延迟,降低被检测为爬虫的概率

内容的提问来源于stack exchange,提问作者David Essien

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 19:37:02