Beautiful Soup提取Booking.com酒店设施不全问题求助
Booking.com酒店设施爬取不全问题排查
近期使用Python的Beautiful Soup爬取Booking.com酒店数据并存入pandas DataFrame时,其余属性均能正确提取,但get_Hotel_Facilities函数仅能返回部分设施,网站上存在的部分设施无法被提取,已尝试更换对应类和属性仍未解决。
提取设施的函数代码
def get_Hotel_Facilities(soup): try: title = soup.find_all("div", attrs={"class":"db29ecfbe2 c21a2f2d97 fe87d598e8"}) new_list = [] # Inner NavigatableString Object for i in range(len(title)): new_list.append(title[i].text.strip()) except AttributeError: new_list="" return new_list
主爬取代码
page_no=0 d = {"Hotel_Name":[],"Hotel_Rating":[],"Room_type":[],"Room_price":[],"Room_sqft":[],"Facilities":[],"Location":[]} while (page_no<=25): URL = f"https://www.booking.com/searchresults.html?aid=304142&label=gen173rf-1FCAEoggI46AdIM1gDaGyIAQGYATG4ARfIAQzYAQHoAQH4AQKIAgGiAg1wcm9qZWN0cHJvLmlvqAIDuAKwwPadBsACAdICJDU0NThkNDAzLTM1OTMtNDRmOC1iZWQ0LTdhOTNjOTJmOWJlONgCBeACAQ&sid=2214b1422694e7b065e28995af4e22d9&sb=1&sb_lp=1&src=theme_landing_index&src_elem=sb&error_url=https%3A%2F%2Fwww.booking.com%2Fhotel%2Findex.html%3Faid%3D304142%26label%3Dgen173rf1FCAEoggI46AdIM1gDaGyIAQGYATG4ARfIAQzYAQHoAQH4AQKIAgGiAg1wcm9qZWN0cHJvLmlvqAIDuAKwwPadBsACAdICJDU0NThkNDAzLTM1OTMtNDRmOC1iZWQ0LTdhOTNjOTJmOWJlONgCBeACAQ%26sid%3D2214b1422694e7b065e28995af4e22d9%26&ss=goa&is_ski_area=0&checkin_year=2023&checkin_month=1&checkin_monthday=13&checkout_year=2023&checkout_month=1&checkout_monthday=14&group_adults=2&group_children=0&no_rooms=1&b_h4u_keep_filters=&from_sf=1&offset{page_no}" new_webpage = requests.get(URL, headers=HEADERS) soup = BeautifulSoup(new_webpage.content,"html.parser") links = soup.find_all("a", attrs={"class":"e13098a59f"}) for link in links: new_webpage = requests.get(link.get('href'), headers=HEADERS) new_soup = BeautifulSoup(new_webpage.content, "html.parser") d["Hotel_Name"].append(get_Hotel_Name(new_soup)) d["Hotel_Rating"].append(get_Hotel_Rating(new_soup)) d["Room_type"].append(get_Room_type(new_soup)) d["Room_price"].append(get_Price(new_soup)) d["Room_sqft"].append(get_Room_Sqft(new_soup)) d["Facilities"].append(get_Hotel_Facilities(new_soup)) d["Location"].append(get_Hotel_Location(new_soup)) page_no += 25
问题原因及解决方案
可能的原因
- 设施分散在不同HTML容器:当前函数仅抓取了固定类名的
div,未提取其他容器中的设施内容 - 动态加载内容缺失:部分设施通过JavaScript动态渲染,
requests获取的静态HTML中不包含这部分数据 - 动态类名失效:Booking.com的类名可能是动态生成的,固定类名仅匹配了部分元素
具体解决步骤
排查页面真实结构
- 打开浏览器开发者工具(F12),定位未被提取的设施对应的HTML标签,记录其容器的类名、标签类型或父容器属性
- 例如:部分设施可能在
li标签中,或嵌套在带有data-testid="property-facilities"属性的父容器内
修改函数覆盖多类容器
调整get_Hotel_Facilities函数,同时抓取多个可能的设施容器,避免遗漏:def get_Hotel_Facilities(soup): facilities = [] # 抓取原类名的设施容器 containers1 = soup.find_all("div", attrs={"class":"db29ecfbe2 c21a2f2d97 fe87d598e8"}) for item in containers1: text = item.text.strip() if text: facilities.append(text) # 新增抓取其他容器(根据实际页面结构修改类名) containers2 = soup.find_all("li", attrs={"class":"a51f9c992e"}) for item in containers2: text = item.text.strip() if text and text not in facilities: facilities.append(text) # 基于父容器定位(更稳定的方式) parent_container = soup.find("div", attrs={"data-testid":"property-facilities"}) if parent_container: extra_items = [item.text.strip() for item in parent_container.find_all("div") if item.text.strip()] facilities.extend([item for item in extra_items if item not in facilities]) return facilities if facilities else []处理动态加载内容
若部分设施为动态加载,改用selenium模拟浏览器渲染页面:from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 替换原requests请求部分 driver = webdriver.Chrome() driver.get(link.get('href')) # 等待设施区域加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "db29ecfbe2")) ) new_soup = BeautifulSoup(driver.page_source, "html.parser") # 后续提取操作不变 driver.quit()优化定位逻辑
避免依赖易变的动态类名,优先使用data-testid等稳定属性,或通过标签层级关系定位元素。
内容的提问来源于stack exchange,提问作者BODDAPALLI KOUSHIK
相关产品推荐
相关产品推荐

