You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests模块爬取TripAdvisor餐厅网址失败,求解决方法

解决TripAdvisor餐厅网址爬取失败的问题

你的问题出在选择器无法匹配到目标元素,且TripAdvisor的餐厅网址并非直接暴露在页面可见链接文本中,而是通过跳转链接或内嵌JSON数据提供。以下是两种可靠的解决方法:

方法一:从页面内嵌JSON数据提取(推荐)

TripAdvisor会将餐厅核心信息(包括网址)内嵌在页面的window.__WEB_CONTEXT__变量中,直接解析这个JSON数据可以稳定获取目标网址:

import requests
from bs4 import BeautifulSoup
import json

link = 'https://www.tripadvisor.com/Restaurant_Review-g293995-d19398253-Reviews-Yauatcha_Riyadh-Riyadh_Riyadh_Province.html'

headers = {
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
    'Accept-Encoding': 'gzip, deflate, br, zstd',
    'Accept-Language': 'en-US,en;q=0.9',
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36',
}
with requests.Session() as s:
    s.headers.update(headers)
    res = s.get(link)
    soup = BeautifulSoup(res.text,"lxml")
    
    # 获取餐厅名称和电话
    title = soup.select_one("[data-test-target='restaurant-detail-info'] h1").get_text(strip=True)
    phone = soup.select_one("[data-test-target='restaurant-detail-info'] a[href^='tel:']").get("href")
    
    # 解析内嵌JSON数据提取网址
    website = ""
    script_tag = soup.find("script", text=lambda t: t and "__WEB_CONTEXT__" in t)
    if script_tag:
        # 提取JSON字符串并解析
        json_content = script_tag.text.split("window.__WEB_CONTEXT__ = ")[1].split(";")[0].strip()
        data = json.loads(json_content)
        website = data.get("pageManifest", {}).get("entityInfo", {}).get("websiteUrl", "")
    
    print((title, phone, website))

方法二:解析TripAdvisor跳转链接

如果内嵌JSON数据无法获取,可以解析页面中的跳转链接,提取真实网址:

import requests
from bs4 import BeautifulSoup
from urllib.parse import unquote, urlparse, parse_qs

link = 'https://www.tripadvisor.com/Restaurant_Review-g293995-d19398253-Reviews-Yauatcha_Riyadh-Riyadh_Riyadh_Province.html'

headers = {
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
    'Accept-Encoding': 'gzip, deflate, br, zstd',
    'Accept-Language': 'en-US,en;q=0.9',
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36',
}
with requests.Session() as s:
    s.headers.update(headers)
    res = s.get(link)
    soup = BeautifulSoup(res.text,"lxml")
    
    # 获取餐厅名称和电话
    title = soup.select_one("[data-test-target='restaurant-detail-info'] h1").get_text(strip=True)
    phone = soup.select_one("[data-test-target='restaurant-detail-info'] a[href^='tel:']").get("href")
    
    # 解析跳转链接提取真实网址
    website = ""
    redirect_a_tag = soup.select_one("[data-test-target='restaurant-detail-info'] a[href^='https://www.tripadvisor.com/LinkRedirect']")
    if redirect_a_tag:
        parsed_url = urlparse(redirect_a_tag.get("href"))
        redirect_param = parse_qs(parsed_url.query).get("redirect", [])
        if redirect_param:
            website = unquote(redirect_param[0])
    
    print((title, phone, website))

原代码失败原因

  1. 选择器兼容性问题::soup-contains是BeautifulSoup专属CSS选择器扩展,lxml解析器对其支持有限,无法准确匹配元素。
  2. 页面结构变化:TripAdvisor的"Website"链接文本可能被图标元素替代,无法通过文本内容匹配到目标链接。

内容的提问来源于stack exchange,提问作者MITHU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 19:33:11