You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取Deliveroo餐厅数据仅获20条结果,如何获取全部近百条数据?

Deliveroo爬取全部餐厅的解决方案

问题原因

直接用requests获取的页面是初始加载内容,Deliveroo采用动态懒加载机制——只有当用户滚动页面时,才会通过API请求加载更多餐厅数据,所以初始HTML仅包含约20条结果。

方案一:调用官方API(高效推荐)

Deliveroo的餐厅数据是通过API接口返回的,直接请求接口可以批量获取所有数据,无需模拟浏览器。

步骤与代码

  1. 从原URL中提取目标区域的geohash(示例中为gcpn7n35zy89)
  2. 循环请求分页API,通过offset参数控制分页,直到返回空数据
import requests
import pandas as pd

# 替换为目标区域的geohash
geohash = "gcpn7n35zy89"
base_api_url = f"https://deliveroo.co.uk/api/consumer/v3/restaurants?geohash={geohash}&sort=distance&limit=30&offset="

# 必要请求头,模拟浏览器访问
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

all_restaurant_names = []
offset = 0

while True:
    # 发送分页请求
    response = requests.get(base_api_url + str(offset), headers=headers)
    response_data = response.json()
    
    # 提取当前页的餐厅名称
    current_page_names = [resto["name"] for resto in response_data.get("restaurants", [])]
    if not current_page_names:
        break  # 无更多数据,停止循环
    
    all_restaurant_names.extend(current_page_names)
    offset += 30  # 偏移量递增,获取下一页

# 输出结果
print(f"共抓取到 {len(all_restaurant_names)} 家餐厅:")
for name in all_restaurant_names:
    print(name)

# 保存到CSV文件
pd.DataFrame({"餐厅名称": all_restaurant_names}).to_csv("deliveroo_restaurants.csv", index=False, encoding="utf-8-sig")

方案二:Selenium模拟浏览器滚动(直观易操作)

如果不想分析API,可以用Selenium模拟浏览器滚动加载,让页面自动加载所有餐厅数据后再提取。

步骤与代码

  1. 安装依赖:pip install selenium,并下载对应浏览器的驱动(如ChromeDriver)
  2. 模拟浏览器访问、处理Cookie弹窗、滚动加载所有内容
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import pandas as pd

target_url = "https://deliveroo.co.uk/restaurants/oxford/port-meadow?geohash=gcpn7n35zy89&sort=distance"

# 初始化Chrome浏览器(需确保驱动路径正确)
driver = webdriver.Chrome()
driver.get(target_url)

# 处理Cookie同意弹窗(根据页面实际按钮调整定位)
try:
    accept_cookie_btn = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Accept')]"))
    )
    accept_cookie_btn.click()
except:
    pass  # 无弹窗则跳过

# 滚动页面加载所有数据
last_page_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # 等待数据加载
    new_page_height = driver.execute_script("return document.body.scrollHeight")
    if new_page_height == last_page_height:
        break  # 页面高度不变,说明已加载完所有内容
    last_page_height = new_page_height

# 提取所有餐厅名称
restaurant_elements = driver.find_elements(By.CSS_SELECTOR, "li.HomeFeedUILines-8bd2eacc5b5ba98e p")
all_names = [elem.text for elem in restaurant_elements if elem.text.strip()]

# 关闭浏览器
driver.quit()

# 输出结果
print(f"共抓取到 {len(all_names)} 家餐厅:")
for name in all_names:
    print(name)

# 保存到CSV文件
pd.DataFrame({"餐厅名称": all_names}).to_csv("deliveroo_restaurants_selenium.csv", index=False, encoding="utf-8-sig")

注意事项

  • API方法需注意请求频率,避免被反爬限制,可适当添加time.sleep(1)控制请求间隔
  • Selenium的Cookie弹窗定位需根据页面实际文案调整,若按钮文本不是"Accept",替换为对应内容即可

内容的提问来源于stack exchange,提问作者AAli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 18:15:40