You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup与requests自动化网页抓取的分页URL生成?

如何自动化Trustpilot分页URL的生成与数据抓取?

当然可以!手动维护分页URL确实挺繁琐的,我们可以通过分析URL结构+循环逻辑完全自动化这个流程,甚至还能自动判断什么时候停止爬取(比如页面没有更多餐厅数据时)。

第一步:分析URL结构

先看你手动整理的URL,核心规律很明显:

  • 基础URL是 https://www.trustpilot.com/categories/restaurants_bars?numberofreviews=0&status=all&timeperiod=0
  • 分页只需要添加 &page=N 参数(N是页码,从1开始,你的第一页其实可以加上&page=1,不影响访问)

第二步:自动化生成URL+循环爬取

我们可以用一个循环来生成页码,同时加入判断逻辑:如果当前页面没有抓取到餐厅数据,就停止循环(避免无效请求)。另外还修正了你原代码里的小错误(比如导入语句语法问题、变量名大小写不一致),加上基础的异常处理(防止请求失败导致程序崩溃)。

修改后的代码示例:

from bs4 import BeautifulSoup
import requests

# 定义基础URL,把分页参数留空,后面循环补充
base_url = "https://www.trustpilot.com/categories/restaurants_bars?numberofreviews=0&status=all&timeperiod=0"
page_num = 1

while True:
    # 生成当前分页的URL
    current_url = f"{base_url}&page={page_num}"
    try:
        # 发送请求,添加超时防止卡住
        response = requests.get(current_url, timeout=10)
        response.raise_for_status()  # 如果请求状态码不是200,抛出异常
        html_text = response.text
        soup = BeautifulSoup(html_text, 'lxml')
        
        # 抓取餐厅列表
        restaurants = soup.find_all('div', class_='categoryBusinessListWrapper___14CgD')
        
        # 如果没有餐厅数据,说明已经到最后一页,退出循环
        if not restaurants:
            print(f"第{page_num}页没有数据,停止爬取")
            break
        
        # 提取数据(保留你原来的逻辑)
        for index, restaurant in enumerate(restaurants):
            tags = restaurant.find_all('a', class_='internal___1jK0Z wrapper___26yB4')
            for tag in tags:
                restaurant_name = tag.find('div', class_='businessTitle___152-c').text.split(',')[0]
                ratings = tag.find('div', class_='textRating___3F1NO')
                location = tag.find('span', class_='locationZipcodeAndCity___33EfU')
                more_info = tag['href']
                
                # 这里可以添加保存数据的逻辑,比如写入CSV/数据库
                print(f"餐厅名称: {restaurant_name}, 评分: {ratings.text if ratings else '无'}, 位置: {location.text if location else '无'}")
        
        # 页码加1,继续下一页
        page_num += 1
        
    except requests.exceptions.RequestException as e:
        print(f"请求第{page_num}页时出错: {e}")
        # 可以选择跳过当前页或者退出循环,这里选择跳过
        page_num += 1
        continue

关键优化点说明

  • 动态URL生成:通过f-string拼接基础URL和页码参数,不用手动写每一页的URL
  • 自动终止逻辑:检查restaurants是否为空,为空则停止爬取,避免无效请求
  • 异常处理:捕获请求过程中的异常(比如超时、404、500),让程序更健壮
  • 代码规范:修正了原代码的导入错误和变量名不一致问题

这样你就不用再手动维护URL列表了,程序会自动从第1页开始爬,直到没有更多数据为止。

内容的提问来源于stack exchange,提问作者Shashank Tripathi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:47:22