You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取HostelWorld巴黎酒店详情链接失败的问题求助

问题:爬取Hostelworld酒店详情链接失败,仅返回['/']

需求与问题描述

用户想要获取巴黎Hostelworld搜索结果页中的所有酒店详情链接,目标页面地址:

https://www.french.hostelworld.com/s?q=Paris,%20Ile-de-France,%20France&country=France&city=Paris&type=city&id=14&from=2021-04-30&to=2021-05-03&guests=2&page=1

期望得到类似如下格式的详情链接列表:

['https://www.french.hostelworld.com/pwa/hosteldetails.php/R-sidence-Internationale-de-Paris/Paris/294403?from=2021-04-30&to=2021-05-03&guests=2', ...]

但运行自己编写的脚本后,仅得到结果:['/'],无法获取到预期的酒店链接。

问题原因分析

  1. 页面动态加载:目标网站属于现代动态网站,酒店列表内容是通过JavaScript异步渲染生成的。直接使用requests.get()只能获取到页面的静态HTML框架,无法拿到JS加载后的实际酒店数据。
  2. 元素定位不准确:原脚本中通过find("div", {"class": "page-inner"}).find_all('a', href=True)定位链接,但静态HTML中这个div下仅包含网站首页的链接/,实际的酒店链接在JS渲染完成后才会出现在页面中。

解决思路与方案

思路1:直接调用网站API(推荐)

大多数旅游预订网站会通过API接口返回数据,你可以通过浏览器开发者工具找到对应的接口:

  • 打开浏览器F12开发者工具,切换到Network标签
  • 刷新目标页面,筛选XHR/Fetch类型的请求,查找包含酒店列表的接口(通常命名包含search、properties等关键词)
  • 复制该接口的请求URL和参数,直接用requests请求并解析JSON数据,提取其中的酒店详情链接。

示例代码(假设找到API接口):

import requests

api_url = "https://www.french.hostelworld.com/api/search/properties"
params = {
    "q": "Paris, Ile-de-France, France",
    "country": "France",
    "city": "Paris",
    "type": "city",
    "id": "14",
    "from": "2021-04-30",
    "to": "2021-05-03",
    "guests": "2",
    "page": "1"
}

headers = {
    # 可能需要添加User-Agent等请求头,模拟浏览器请求
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

response = requests.get(api_url, params=params, headers=headers)
data = response.json()

# 从返回的JSON中提取酒店详情链接,具体字段根据实际API返回调整
hotel_links = []
for property in data.get("properties", []):
    link = property.get("url")
    if link:
        # 拼接完整URL(如果返回的是相对路径)
        if not link.startswith("http"):
            link = "https://www.french.hostelworld.com" + link
        hotel_links.append(link)

print(hotel_links)

思路2:使用动态渲染工具(Selenium/Playwright)

如果找不到API接口,可以用模拟浏览器的工具等待页面JS渲染完成后再提取元素:

示例代码(Selenium)

首先安装依赖:

pip install selenium

同时需要下载对应浏览器的驱动(如ChromeDriver),并确保驱动路径正确:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = 'https://www.french.hostelworld.com/s?q=Paris,%20Ile-de-France,%20France&country=France&city=Paris&type=city&id=14&from=2021-04-30&to=2021-05-03&guests=2&page=1'

# 初始化Chrome浏览器驱动
driver = webdriver.Chrome()
driver.get(url)

# 等待酒店卡片加载完成(根据页面实际元素调整选择器)
wait = WebDriverWait(driver, 10)
hotel_cards = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.property-card")))

# 提取所有酒店详情链接
base_url = 'https://www.french.hostelworld.com'
hotel_links = []
for card in hotel_cards:
    a_tag = card.find_element(By.TAG_NAME, 'a')
    href = a_tag.get_attribute('href')
    if href and href != '/':
        if not href.startswith('http'):
            href = base_url + href
        hotel_links.append(href)

print(hotel_links)

# 关闭浏览器
driver.quit()

总结

原脚本失败的核心原因是无法获取动态渲染的内容,优先推荐使用API接口的方式,效率更高且更稳定;如果API接口难以找到,再使用Selenium等工具模拟浏览器渲染页面来提取链接。

内容的提问来源于stack exchange,提问作者FalconMelee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 15:42:34