使用BeautifulSoup提取HostelWorld巴黎酒店详情链接失败的问题求助
问题:爬取Hostelworld酒店详情链接失败,仅返回
['/'] 需求与问题描述
用户想要获取巴黎Hostelworld搜索结果页中的所有酒店详情链接,目标页面地址:
https://www.french.hostelworld.com/s?q=Paris,%20Ile-de-France,%20France&country=France&city=Paris&type=city&id=14&from=2021-04-30&to=2021-05-03&guests=2&page=1
期望得到类似如下格式的详情链接列表:
['https://www.french.hostelworld.com/pwa/hosteldetails.php/R-sidence-Internationale-de-Paris/Paris/294403?from=2021-04-30&to=2021-05-03&guests=2', ...]
但运行自己编写的脚本后,仅得到结果:['/'],无法获取到预期的酒店链接。
问题原因分析
- 页面动态加载:目标网站属于现代动态网站,酒店列表内容是通过JavaScript异步渲染生成的。直接使用
requests.get()只能获取到页面的静态HTML框架,无法拿到JS加载后的实际酒店数据。 - 元素定位不准确:原脚本中通过
find("div", {"class": "page-inner"}).find_all('a', href=True)定位链接,但静态HTML中这个div下仅包含网站首页的链接/,实际的酒店链接在JS渲染完成后才会出现在页面中。
解决思路与方案
思路1:直接调用网站API(推荐)
大多数旅游预订网站会通过API接口返回数据,你可以通过浏览器开发者工具找到对应的接口:
- 打开浏览器F12开发者工具,切换到Network标签
- 刷新目标页面,筛选XHR/Fetch类型的请求,查找包含酒店列表的接口(通常命名包含
search、properties等关键词) - 复制该接口的请求URL和参数,直接用
requests请求并解析JSON数据,提取其中的酒店详情链接。
示例代码(假设找到API接口):
import requests api_url = "https://www.french.hostelworld.com/api/search/properties" params = { "q": "Paris, Ile-de-France, France", "country": "France", "city": "Paris", "type": "city", "id": "14", "from": "2021-04-30", "to": "2021-05-03", "guests": "2", "page": "1" } headers = { # 可能需要添加User-Agent等请求头,模拟浏览器请求 "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } response = requests.get(api_url, params=params, headers=headers) data = response.json() # 从返回的JSON中提取酒店详情链接,具体字段根据实际API返回调整 hotel_links = [] for property in data.get("properties", []): link = property.get("url") if link: # 拼接完整URL(如果返回的是相对路径) if not link.startswith("http"): link = "https://www.french.hostelworld.com" + link hotel_links.append(link) print(hotel_links)
思路2:使用动态渲染工具(Selenium/Playwright)
如果找不到API接口,可以用模拟浏览器的工具等待页面JS渲染完成后再提取元素:
示例代码(Selenium)
首先安装依赖:
pip install selenium
同时需要下载对应浏览器的驱动(如ChromeDriver),并确保驱动路径正确:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = 'https://www.french.hostelworld.com/s?q=Paris,%20Ile-de-France,%20France&country=France&city=Paris&type=city&id=14&from=2021-04-30&to=2021-05-03&guests=2&page=1' # 初始化Chrome浏览器驱动 driver = webdriver.Chrome() driver.get(url) # 等待酒店卡片加载完成(根据页面实际元素调整选择器) wait = WebDriverWait(driver, 10) hotel_cards = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.property-card"))) # 提取所有酒店详情链接 base_url = 'https://www.french.hostelworld.com' hotel_links = [] for card in hotel_cards: a_tag = card.find_element(By.TAG_NAME, 'a') href = a_tag.get_attribute('href') if href and href != '/': if not href.startswith('http'): href = base_url + href hotel_links.append(href) print(hotel_links) # 关闭浏览器 driver.quit()
总结
原脚本失败的核心原因是无法获取动态渲染的内容,优先推荐使用API接口的方式,效率更高且更稳定;如果API接口难以找到,再使用Selenium等工具模拟浏览器渲染页面来提取链接。
内容的提问来源于stack exchange,提问作者FalconMelee
相关产品推荐
相关产品推荐

