You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Beautiful Soup提取Booking.com酒店设施不全问题求助

Booking.com酒店设施爬取不全问题排查

近期使用Python的Beautiful Soup爬取Booking.com酒店数据并存入pandas DataFrame时,其余属性均能正确提取,但get_Hotel_Facilities函数仅能返回部分设施,网站上存在的部分设施无法被提取,已尝试更换对应类和属性仍未解决。

提取设施的函数代码

def get_Hotel_Facilities(soup):
    try:
        title = soup.find_all("div", attrs={"class":"db29ecfbe2 c21a2f2d97 fe87d598e8"})
        new_list = []
        # Inner NavigatableString Object
        for i in range(len(title)):
          new_list.append(title[i].text.strip())

    except AttributeError:
       new_list=""

    return new_list

主爬取代码

page_no=0
d = {"Hotel_Name":[],"Hotel_Rating":[],"Room_type":[],"Room_price":[],"Room_sqft":[],"Facilities":[],"Location":[]}
while (page_no<=25):
     URL = f"https://www.booking.com/searchresults.html?aid=304142&label=gen173rf-1FCAEoggI46AdIM1gDaGyIAQGYATG4ARfIAQzYAQHoAQH4AQKIAgGiAg1wcm9qZWN0cHJvLmlvqAIDuAKwwPadBsACAdICJDU0NThkNDAzLTM1OTMtNDRmOC1iZWQ0LTdhOTNjOTJmOWJlONgCBeACAQ&sid=2214b1422694e7b065e28995af4e22d9&sb=1&sb_lp=1&src=theme_landing_index&src_elem=sb&error_url=https%3A%2F%2Fwww.booking.com%2Fhotel%2Findex.html%3Faid%3D304142%26label%3Dgen173rf1FCAEoggI46AdIM1gDaGyIAQGYATG4ARfIAQzYAQHoAQH4AQKIAgGiAg1wcm9qZWN0cHJvLmlvqAIDuAKwwPadBsACAdICJDU0NThkNDAzLTM1OTMtNDRmOC1iZWQ0LTdhOTNjOTJmOWJlONgCBeACAQ%26sid%3D2214b1422694e7b065e28995af4e22d9%26&ss=goa&is_ski_area=0&checkin_year=2023&checkin_month=1&checkin_monthday=13&checkout_year=2023&checkout_month=1&checkout_monthday=14&group_adults=2&group_children=0&no_rooms=1&b_h4u_keep_filters=&from_sf=1&offset{page_no}"
     new_webpage = requests.get(URL, headers=HEADERS)
     soup = BeautifulSoup(new_webpage.content,"html.parser")
     links = soup.find_all("a", attrs={"class":"e13098a59f"})
     for link in links:
        new_webpage = requests.get(link.get('href'), headers=HEADERS)
        new_soup = BeautifulSoup(new_webpage.content, "html.parser")
        d["Hotel_Name"].append(get_Hotel_Name(new_soup))
        d["Hotel_Rating"].append(get_Hotel_Rating(new_soup))
        d["Room_type"].append(get_Room_type(new_soup))
        d["Room_price"].append(get_Price(new_soup))
        d["Room_sqft"].append(get_Room_Sqft(new_soup))
        d["Facilities"].append(get_Hotel_Facilities(new_soup))
        d["Location"].append(get_Hotel_Location(new_soup))

     page_no += 25

问题原因及解决方案

可能的原因

  1. 设施分散在不同HTML容器:当前函数仅抓取了固定类名的div,未提取其他容器中的设施内容
  2. 动态加载内容缺失:部分设施通过JavaScript动态渲染,requests获取的静态HTML中不包含这部分数据
  3. 动态类名失效:Booking.com的类名可能是动态生成的,固定类名仅匹配了部分元素

具体解决步骤

  1. 排查页面真实结构

    • 打开浏览器开发者工具(F12),定位未被提取的设施对应的HTML标签,记录其容器的类名、标签类型或父容器属性
    • 例如:部分设施可能在li标签中,或嵌套在带有data-testid="property-facilities"属性的父容器内
  2. 修改函数覆盖多类容器
    调整get_Hotel_Facilities函数,同时抓取多个可能的设施容器,避免遗漏:

    def get_Hotel_Facilities(soup):
        facilities = []
        # 抓取原类名的设施容器
        containers1 = soup.find_all("div", attrs={"class":"db29ecfbe2 c21a2f2d97 fe87d598e8"})
        for item in containers1:
            text = item.text.strip()
            if text:
                facilities.append(text)
        # 新增抓取其他容器(根据实际页面结构修改类名)
        containers2 = soup.find_all("li", attrs={"class":"a51f9c992e"})
        for item in containers2:
            text = item.text.strip()
            if text and text not in facilities:
                facilities.append(text)
        # 基于父容器定位(更稳定的方式)
        parent_container = soup.find("div", attrs={"data-testid":"property-facilities"})
        if parent_container:
            extra_items = [item.text.strip() for item in parent_container.find_all("div") if item.text.strip()]
            facilities.extend([item for item in extra_items if item not in facilities])
        return facilities if facilities else []
    
  3. 处理动态加载内容
    若部分设施为动态加载,改用selenium模拟浏览器渲染页面:

    from selenium import webdriver
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    # 替换原requests请求部分
    driver = webdriver.Chrome()
    driver.get(link.get('href'))
    # 等待设施区域加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "db29ecfbe2"))
    )
    new_soup = BeautifulSoup(driver.page_source, "html.parser")
    # 后续提取操作不变
    driver.quit()
    
  4. 优化定位逻辑
    避免依赖易变的动态类名,优先使用data-testid等稳定属性,或通过标签层级关系定位元素。


内容的提问来源于stack exchange,提问作者BODDAPALLI KOUSHIK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 19:30:37