Selenium及BeautifulSoup爬取Zillow房源链接无法返回全部元素如何解决
Zillow出租房源爬取不全问题解决方案
问题根因
- 静态请求(BeautifulSoup方案):Zillow房源采用动态异步加载逻辑,首次返回的HTML仅包含首屏前9条左右数据,剩余房源需用户滚动页面触发后端接口请求后才会渲染到DOM中,直接通过requests请求页面只能拿到首屏少量数据。
- Selenium原生滚动方案:原有逻辑直接将滚动容器拉到底部,无法触发Zillow懒加载的视口检测规则,中间区间的房源不会加载;同时未处理分页逻辑,单页加载到40条左右上限后没有跳转下一页,且Selenium的自动化特征容易被Zillow反爬系统识别,限制返回数据量。
解决方案
方案1:优化Selenium爬取逻辑
首先添加反爬配置隐藏自动化特征,再修改为分段滚动逻辑,同时适配分页跳转:
import time from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 配置反爬参数 chrome_options = Options() chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option("useAutomationExtension", False) driver = webdriver.Chrome(options=chrome_options) # 隐藏webdriver标识 driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", { "source": "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})" }) ZILLOW_URL = "你的Zillow搜索链接" driver.get(ZILLOW_URL) time.sleep(3) scroll_container = driver.find_element(By.ID, 'search-page-list-container') last_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container) all_links = set() # 自动去重 while True: # 分段滚动半屏,触发懒加载 driver.execute_script("arguments[0].scrollTop += arguments[0].offsetHeight/2", scroll_container) time.sleep(2) # 提取当前已加载的房源链接 link_tags = driver.find_elements(By.CSS_SELECTOR, ".list-card-info a") for tag in link_tags: href = tag.get_attribute("href") if href: all_links.add(href) # 判断是否滚动到当前页底部 new_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container) if new_height == last_height: # 尝试点击下一页 try: next_btn = WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'a[title="Next page"]'))) next_btn.click() time.sleep(3) last_height = driver.execute_script("return arguments[0].scrollHeight", scroll_container) except: # 没有下一页则退出循环 break last_height = new_height # 输出所有链接 print(list(all_links)) driver.quit()
方案2:直接调用后端接口爬取(效率更高)
打开浏览器F12开发者工具,切换到「网络」选项卡,筛选「Fetch/XHR」请求,滚动页面时可以找到搜索接口,请求返回为JSON格式,直接包含全量房源数据,无需解析HTML。接口请求参数中的pagination字段可控制分页,修改页码即可批量拉取所有房源。
注意:请求接口时需要完整复制浏览器中的请求头(包含cookie、user-agent、referer等字段),控制请求频率,必要时使用代理IP池避免被反爬封禁。
内容的提问来源于stack exchange,提问作者Tharu
相关产品推荐
相关产品推荐

