You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无限滚动爬取问题:Selenium+BeautifulSoup仅获40条数据(共440条)

解决Airbnb滚动爬取仅获取前40条数据的问题

你的问题核心是滚动加载逻辑未正确触发Airbnb的动态数据加载,导致页面只加载了初始的40条房源。以下是具体修正方案:

问题分析

  1. 原滚动逻辑使用Keys.END直接跳转到页面底部,这种方式可能不会触发Airbnb的滚动加载机制(Airbnb需要滚动到接近底部时才会请求新数据)。
  2. 等待条件EC.presence_of_all_elements_located((By.CLASS_NAME, 'dir-ltr'))过于宽泛,dir-ltr是页面通用类,无法准确判断新数据是否加载完成。
  3. 滚动终止条件判断不准确,动态加载后页面scrollHeight会更新,原条件可能提前终止循环。

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.keys import Keys
import time
import json
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from bs4 import BeautifulSoup

# 初始化驱动
path_to_driver = Service(executable_path=r'D:/Projects/web_scraping/Chrome_driver/chromedriver.exe')
driver = webdriver.Chrome(service=path_to_driver)

url = "https://www.airbnb.com/?tab_id=home_tab&refinement_paths%5B%5D=%2Fhomes&search_mode=flex_destinations_search&flexible_trip_lengths%5B%5D=one_week&location_search=MIN_MAP_BOUNDS&monthly_start_date=2023-09-01&monthly_length=3&price_filter_input_type=0&price_filter_num_nights=5&channel=EXPLORE&search_type=category_change&category_tag=Tag%3A8678"
driver.get(url)
driver.maximize_window()
wait = WebDriverWait(driver, 30)

# 处理Cookie弹窗(Airbnb会强制要求,不处理可能影响交互)
try:
    wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[data-testid="accept-btn"]'))).click()
except:
    pass  # 若没有弹窗则跳过

def scroll_to_end():
    last_height = driver.execute_script("return document.body.scrollHeight")
    while True:
        # 滚动到当前页面底部(模拟用户滚动行为)
        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        # 等待数据加载,可根据网络情况调整时间
        time.sleep(3)
        # 等待加载动画消失(Airbnb加载时会显示 spinner)
        try:
            wait.until(EC.invisibility_of_element_located((By.CLASS_NAME, "_1hxyyw3")))
        except:
            pass
        # 获取新的页面高度
        new_height = driver.execute_script("return document.body.scrollHeight")
        # 若高度不再变化,说明所有数据加载完成
        if new_height == last_height:
            break
        last_height = new_height

scroll_to_end()

# 提取数据
html = driver.page_source
soup = BeautifulSoup(html, 'lxml')
data = soup.select_one('#data-deferred-state')
data = json.loads(data.text)
items = data["niobeMinimalClientData"][0][1]["data"]["presentation"]["explore"]["sections"]["sectionIndependentData"]["staysMapSearch"]["mapSearchResults"]

for item in items:
    item_id = item['listing']['id']
    print(item_id)

driver.quit()

关键优化点

  • 使用window.scrollTo JavaScript滚动,更贴近用户滚动行为,确保触发Airbnb的加载机制。
  • 通过对比滚动前后的页面高度,准确判断是否所有数据加载完成。
  • 增加Cookie弹窗处理,避免页面交互被阻塞。
  • 等待加载动画消失,确保新数据完全渲染后再继续滚动。

内容的提问来源于stack exchange,提问作者Md. Zahirul Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:44:53