如何用Scrapy抓取动态加载网站?含charliesmithrealty.net爬取需求
解决Scrapy结合Selenium抓取动态过滤搜索结果的问题
我仔细看了你提供的代码和遇到的问题,咱们一步步拆解问题并给出可行的修复方案:
你的代码里的核心问题
- 过时的元素定位方法:
find_element_by_xpath这类旧方法已经被Selenium弃用,必须改用find_element(By.XPATH, "xxx")的标准格式 - Driver生命周期错误:你在循环完分页后先关闭了driver,再去获取
page_source,这时候driver已经销毁,必然会抛出异常 - 分页逻辑漏洞:你先一直点击下一页直到无法点击,这时候之前页面的房源数据已经被新页面覆盖,最终只能拿到最后一页的内容
- 元素定位不够精准:部分XPath选择器没有考虑页面动态渲染的延迟,容易出现元素未找到的报错
修复后的完整代码
下面是调整后的代码,解决了上述所有问题,能稳定抓取每一页的搜索结果:
import scrapy from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from scrapy.selector import Selector class RealSpider(scrapy.Spider): name = 'real' start_urls = ['https://charliesmithrealty.net'] def __init__(self): # 初始化Chrome Driver,添加无头模式提升运行效率 options = webdriver.ChromeOptions() options.add_argument('--headless=new') options.add_argument('--window-size=1920,1080') self.driver = webdriver.Chrome(options=options) self.wait = WebDriverWait(self.driver, 10) # 显式等待替代time.sleep,更可靠 def parse(self, response): self.driver.get(response.url) # 1. 输入并选择邮政编码 location_input = self.wait.until( EC.presence_of_element_located((By.ID, 'bfg-location-input')) ) location_input.send_keys("80023") # 等待下拉选项出现并点击匹配的邮政编码 zip_option = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//li[@lookup_field='zip_code' and contains(text(), '80023')]")) ) zip_option.click() # 2. 设置价格区间 # 点击最小价格下拉框 min_price_dropdown = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'bfg-dropdown-min')]//button")) ) min_price_dropdown.click() min_price_option = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//li[contains(@class, 'bfg-option-list-min')]//li[text()='$150K']")) ) min_price_option.click() # 点击最大价格下拉框 max_price_dropdown = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'bfg-dropdown-max')]//button")) ) max_price_dropdown.click() max_price_option = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//li[contains(@class, 'bfg-option-list-max')]//li[text()='$600K']")) ) max_price_option.click() # 3. 设置卧室和卫生间数量 # 卧室下拉选择(自定义下拉框需用JS触发事件) beds_select = self.wait.until( EC.presence_of_element_located((By.ID, 'BedsDropdown')) ) self.driver.execute_script("arguments[0].value = '4';", beds_select) self.driver.execute_script("arguments[0].dispatchEvent(new Event('change'));", beds_select) # 卫生间下拉选择 baths_select = self.wait.until( EC.presence_of_element_located((By.ID, 'BathsDropdown')) ) self.driver.execute_script("arguments[0].value = '3';", baths_select) self.driver.execute_script("arguments[0].dispatchEvent(new Event('change'));", baths_select) # 4. 执行搜索 search_btn = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//button[@aria-label='Submit Search']")) ) search_btn.click() self.wait.until( EC.presence_of_element_located((By.ID, 'bfg-map-gallery')) ) # 等待搜索结果加载完成 # 5. 分页抓取每一页数据 while True: # 抓取当前页的房源URL html = self.driver.page_source resp = Selector(text=html) for item in resp.xpath("//div[@id='bfg-map-gallery']//mbb-galleryitem/div/a"): yield { 'property_url': item.xpath("./@href").get() } # 尝试点击下一页 try: next_btn = self.wait.until( EC.element_to_be_clickable((By.XPATH, "//a[@data-page='next' and not(@disabled)]")) ) next_btn.click() # 等待新页面加载完成(通过判断旧元素失效实现) self.wait.until( EC.staleness_of(resp.xpath("//div[@id='bfg-map-gallery']").get()) ) except: # 没有下一页时退出循环 break # 最后关闭Driver self.driver.quit()
关键优化点说明
- 用显式等待替代time.sleep:
WebDriverWait会等待元素出现/可点击后再执行操作,避免因网络延迟导致的元素未加载问题,比固定等待时间更灵活可靠 - 精准定位元素:优先用
ID定位,XPath选择器添加更具体的条件,避免因页面结构小变化导致定位失败 - 处理自定义下拉框:部分下拉框是前端自定义组件,直接设置value后需要触发
change事件,才能让页面识别到选择操作 - 分页实时抓取:每点击一次下一页就立即抓取当前页内容,避免数据被新页面覆盖
- 无头模式运行:减少资源占用,更适合服务器或后台运行爬虫
关于API接口的补充建议
你提到没找到JSON接口,这类房产网站通常用React/Vue等前端框架渲染,数据可能通过GraphQL或加密接口传输。如果不想用Selenium,可以尝试:
- 查看Network面板的Fetch/XHR标签,筛选
graphql相关请求,很多这类网站会用GraphQL查询数据 - 复制请求Headers中的
authorization或cookie等认证信息,直接构造请求获取数据
不过如果API接口有反爬机制,Selenium仍然是最直接、最不易被拦截的解决方案。
内容的提问来源于stack exchange,提问作者Raisul Islam
相关产品推荐
相关产品推荐

