You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的Beautiful Soup处理Cookie并提取餐厅名称

解决TripAdvisor Cookie限制的爬虫方案

方法1:手动注入浏览器Cookie

TripAdvisor会校验会话Cookie,你可以先在浏览器打开目标页面,通过开发者工具(F12)的「网络」面板复制完整的Cookie字符串,再添加到请求头中:

import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36",
    "Cookie": "替换为你从浏览器复制的完整Cookie字符串"
}

url = "https://www.tripadvisor.co.uk/FindRestaurants?geo=155019&diets=10697&establishmentTypes=10591&minimumTravelerRating=TRAVELER_RATING_MEDIUM&broadened=false"

response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, "html.parser")

# 提取餐厅名称示例
restaurant_names = soup.find_all("div", class_="result-title")
for name in restaurant_names:
    print(name.get_text(strip=True))

方法2:用requests.Session维持会话

Session会自动管理Cookie生命周期,先访问TripAdvisor首页获取初始会话Cookie,再请求目标页面:

import requests
from bs4 import BeautifulSoup

session = requests.Session()
session.headers.update({
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36"
})

# 先访问首页获取初始Cookie
session.get("https://www.tripadvisor.co.uk")

url = "https://www.tripadvisor.co.uk/FindRestaurants?geo=155019&diets=10697&establishmentTypes=10591&minimumTravelerRating=TRAVELER_RATING_MEDIUM&broadened=false"
response = session.get(url)
soup = BeautifulSoup(response.content, "html.parser")

# 提取餐厅名称示例
restaurant_names = soup.find_all("a", class_="Lwqic Cj b")
for name in restaurant_names:
    print(name.get_text(strip=True))

方法3:Selenium模拟浏览器(推荐)

如果网站存在动态渲染或更严格的反爬机制,Selenium会自动处理Cookie加载和JS渲染,直接获取完整页面内容:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

# 配置Chrome无头模式(可选,隐藏浏览器窗口)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
url = "https://www.tripadvisor.co.uk/FindRestaurants?geo=155019&diets=10697&establishmentTypes=10591&minimumTravelerRating=TRAVELER_RATING_MEDIUM&broadened=false"

driver.get(url)
# 隐式等待10秒,确保页面元素加载完成
driver.implicitly_wait(10)

# 获取完整页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")

# 提取餐厅名称示例
restaurant_names = soup.find_all("a", class_="Lwqic Cj b")
for name in restaurant_names:
    print(name.get_text(strip=True))

driver.quit()

注意事项:

  • 方法1的Cookie会过期,需要定期从浏览器重新复制
  • 方法2仍可能触发反爬,可添加随机请求延迟、轮换User-Agent
  • 方法3需要提前安装Selenium库和对应浏览器驱动(如ChromeDriver)

内容的提问来源于stack exchange,提问作者R Sandy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 13:13:10