You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python BeautifulSoup访问quotes.toscrape.com下一页及后续页面?

批量爬取quotes.toscrape.com多页面内容的解决方案

你的猜测是对的——这个网站的下一页按钮href是相对路径(比如/page/2/),直接请求会因为URL不完整报错。下面是完整的爬取方案,包含相对路径处理和循环爬取逻辑:

核心解决思路

  1. 用urllib.parse.urljoin将基础URL和相对路径拼接成完整请求地址,避免手动拼接出错
  2. 循环检测下一页按钮,直到没有下一页时终止爬取
  3. 每一页请求成功后提取目标数据(名言、作者、标签)

完整代码示例

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

# 网站基础URL
base_url = "http://quotes.toscrape.com"
current_page_url = base_url

while True:
    # 发送HTTP请求,若请求失败直接抛出异常
    response = requests.get(current_page_url)
    response.raise_for_status()
    
    # 解析页面内容
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 提取当前页面的名言数据
    quote_items = soup.find_all("div", class_="quote")
    for item in quote_items:
        quote_text = item.find("span", class_="text").get_text()
        author_name = item.find("small", class_="author").get_text()
        tag_list = [tag.get_text() for tag in item.find_all("a", class_="tag")]
        
        # 打印或保存数据,这里以打印为例
        print(f"名言:{quote_text}")
        print(f"作者:{author_name}")
        print(f"标签:{', '.join(tag_list)}\n---")
    
    # 查找下一页按钮
    next_page_element = soup.find("li", class_="next")
    if not next_page_element:
        # 没有下一页,终止循环
        break
    
    # 拼接下一页完整URL
    next_page_href = next_page_element.find("a")["href"]
    current_page_url = urljoin(base_url, next_page_href)

print("所有页面爬取完成")

关键细节说明

  • urljoin会自动处理相对路径的拼接,比如把http://quotes.toscrape.com和/page/2/拼成http://quotes.toscrape.com/page/2/
  • response.raise_for_status()能快速定位请求失败的问题(比如网络错误、404页面)
  • 下一页的判断依赖li.next元素,这个选择器是网站固定的,若网站结构变更需要同步调整

内容的提问来源于stack exchange,提问作者Dylan Scully

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 22:39:55