无限滚动爬取问题:Selenium+BeautifulSoup仅获40条数据(共440条)
解决Airbnb滚动爬取仅获取前40条数据的问题
你的问题核心是滚动加载逻辑未正确触发Airbnb的动态数据加载,导致页面只加载了初始的40条房源。以下是具体修正方案:
问题分析
- 原滚动逻辑使用
Keys.END直接跳转到页面底部,这种方式可能不会触发Airbnb的滚动加载机制(Airbnb需要滚动到接近底部时才会请求新数据)。 - 等待条件
EC.presence_of_all_elements_located((By.CLASS_NAME, 'dir-ltr'))过于宽泛,dir-ltr是页面通用类,无法准确判断新数据是否加载完成。 - 滚动终止条件判断不准确,动态加载后页面
scrollHeight会更新,原条件可能提前终止循环。
修正后的代码
from selenium import webdriver from selenium.webdriver.common.keys import Keys import time import json from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait from bs4 import BeautifulSoup # 初始化驱动 path_to_driver = Service(executable_path=r'D:/Projects/web_scraping/Chrome_driver/chromedriver.exe') driver = webdriver.Chrome(service=path_to_driver) url = "https://www.airbnb.com/?tab_id=home_tab&refinement_paths%5B%5D=%2Fhomes&search_mode=flex_destinations_search&flexible_trip_lengths%5B%5D=one_week&location_search=MIN_MAP_BOUNDS&monthly_start_date=2023-09-01&monthly_length=3&price_filter_input_type=0&price_filter_num_nights=5&channel=EXPLORE&search_type=category_change&category_tag=Tag%3A8678" driver.get(url) driver.maximize_window() wait = WebDriverWait(driver, 30) # 处理Cookie弹窗(Airbnb会强制要求,不处理可能影响交互) try: wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[data-testid="accept-btn"]'))).click() except: pass # 若没有弹窗则跳过 def scroll_to_end(): last_height = driver.execute_script("return document.body.scrollHeight") while True: # 滚动到当前页面底部(模拟用户滚动行为) driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待数据加载,可根据网络情况调整时间 time.sleep(3) # 等待加载动画消失(Airbnb加载时会显示 spinner) try: wait.until(EC.invisibility_of_element_located((By.CLASS_NAME, "_1hxyyw3"))) except: pass # 获取新的页面高度 new_height = driver.execute_script("return document.body.scrollHeight") # 若高度不再变化,说明所有数据加载完成 if new_height == last_height: break last_height = new_height scroll_to_end() # 提取数据 html = driver.page_source soup = BeautifulSoup(html, 'lxml') data = soup.select_one('#data-deferred-state') data = json.loads(data.text) items = data["niobeMinimalClientData"][0][1]["data"]["presentation"]["explore"]["sections"]["sectionIndependentData"]["staysMapSearch"]["mapSearchResults"] for item in items: item_id = item['listing']['id'] print(item_id) driver.quit()
关键优化点
- 使用
window.scrollToJavaScript滚动,更贴近用户滚动行为,确保触发Airbnb的加载机制。 - 通过对比滚动前后的页面高度,准确判断是否所有数据加载完成。
- 增加Cookie弹窗处理,避免页面交互被阻塞。
- 等待加载动画消失,确保新数据完全渲染后再继续滚动。
内容的提问来源于stack exchange,提问作者Md. Zahirul Islam
相关产品推荐
相关产品推荐

