如何用Python的Beautiful Soup处理Cookie并提取餐厅名称
方法1:手动注入浏览器Cookie
TripAdvisor会校验会话Cookie,你可以先在浏览器打开目标页面,通过开发者工具(F12)的「网络」面板复制完整的Cookie字符串,再添加到请求头中:
import requests from bs4 import BeautifulSoup headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36", "Cookie": "替换为你从浏览器复制的完整Cookie字符串" } url = "https://www.tripadvisor.co.uk/FindRestaurants?geo=155019&diets=10697&establishmentTypes=10591&minimumTravelerRating=TRAVELER_RATING_MEDIUM&broadened=false" response = requests.get(url, headers=headers) soup = BeautifulSoup(response.content, "html.parser") # 提取餐厅名称示例 restaurant_names = soup.find_all("div", class_="result-title") for name in restaurant_names: print(name.get_text(strip=True))
方法2:用requests.Session维持会话
Session会自动管理Cookie生命周期,先访问TripAdvisor首页获取初始会话Cookie,再请求目标页面:
import requests from bs4 import BeautifulSoup session = requests.Session() session.headers.update({ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36" }) # 先访问首页获取初始Cookie session.get("https://www.tripadvisor.co.uk") url = "https://www.tripadvisor.co.uk/FindRestaurants?geo=155019&diets=10697&establishmentTypes=10591&minimumTravelerRating=TRAVELER_RATING_MEDIUM&broadened=false" response = session.get(url) soup = BeautifulSoup(response.content, "html.parser") # 提取餐厅名称示例 restaurant_names = soup.find_all("a", class_="Lwqic Cj b") for name in restaurant_names: print(name.get_text(strip=True))
方法3:Selenium模拟浏览器(推荐)
如果网站存在动态渲染或更严格的反爬机制,Selenium会自动处理Cookie加载和JS渲染,直接获取完整页面内容:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup # 配置Chrome无头模式(可选,隐藏浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) url = "https://www.tripadvisor.co.uk/FindRestaurants?geo=155019&diets=10697&establishmentTypes=10591&minimumTravelerRating=TRAVELER_RATING_MEDIUM&broadened=false" driver.get(url) # 隐式等待10秒,确保页面元素加载完成 driver.implicitly_wait(10) # 获取完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # 提取餐厅名称示例 restaurant_names = soup.find_all("a", class_="Lwqic Cj b") for name in restaurant_names: print(name.get_text(strip=True)) driver.quit()
注意事项:
- 方法1的Cookie会过期,需要定期从浏览器重新复制
- 方法2仍可能触发反爬,可添加随机请求延迟、轮换User-Agent
- 方法3需要提前安装Selenium库和对应浏览器驱动(如ChromeDriver)
内容的提问来源于stack exchange,提问作者R Sandy
相关产品推荐
相关产品推荐

