You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+BeautifulSoup爬取TripAdvisor时因分页无法获取全部餐厅链接

解决TripAdvisor马德里餐厅分页爬取问题

你的问题出在仅请求了第一页数据,未处理TripAdvisor的分页逻辑。TripAdvisor通过-oaXX参数控制分页偏移量,XX代表跳过的条目数(每页固定30条,因此第二页偏移量为30,第三页为60,以此类推)。

以下是修改后的代码,支持自动遍历所有分页并提取全部餐厅链接:

import requests
from bs4 import BeautifulSoup
import time

city_name = 'Madrid'
geo_code = '187514'
base_url = "https://www.tripadvisor.com/Restaurants-g{}-{}-oa{}.html"
headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"
}

# 打开文件保存链接,自动管理资源
with open('madrid_restaurants.txt', 'w', encoding='utf-8') as f:
    offset = 0
    while True:
        # 构造当前分页的请求URL
        url = base_url.format(geo_code, city_name, offset)
        response = requests.get(url, headers=headers)
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # 用CSS选择器提取餐厅链接,比正则更稳定
        restaurant_links = soup.select('a[data-test="restaurant-name"]')
        if not restaurant_links:
            # 无更多链接,终止循环
            break
        
        # 遍历提取到的链接,补全完整URL并保存
        for link in restaurant_links:
            full_link = "https://www.tripadvisor.com" + link['href']
            print(full_link)
            f.write(f"{full_link}\n")
        
        # 检查是否存在下一页按钮
        next_page_btn = soup.select_one('a.nav.next')
        if not next_page_btn:
            break
        
        # 更新偏移量,准备请求下一页
        offset += 30
        # 添加延时,避免请求过于频繁触发反爬
        time.sleep(2)

核心优化说明:

  • 分页循环控制:通过offset变量递增30,实现分页URL的动态生成
  • 稳定链接提取:使用BeautifulSoup的CSS选择器替代正则表达式,适配页面结构变化的能力更强
  • 循环终止条件:当无法提取到餐厅链接或无下一页按钮时,自动停止爬取
  • 反爬防护:加入time.sleep(2)降低请求频率,减少被网站封禁的风险
  • 文件操作优化:采用with语句自动管理文件资源,避免遗漏关闭操作

内容的提问来源于stack exchange,提问作者MDON

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:50:30