如何用scrapy-selenium爬取需点击切换Tab才能加载的页面数据
问题根因
返回空值有3个核心原因:
- 通过Selenium点击「公司简介」tab后,页面数据是异步加载的,没有等待数据渲染完成就进行提取
- Scrapy自带的
response对象存储的是页面初始加载完成的静态内容,在Selenium中操作DOM后的新内容不会自动同步到这个response里,直接用原有response.xpath提取无法拿到动态渲染的公司数据 - 你在提取数据前就调用了
driver.quit()关掉了浏览器实例,后续也无法从driver侧拿到最新页面内容
修复方案
首先导入Selenium显式等待依赖包,再修改crawl_mainpage方法即可:
# 新增导入放在文件头部 from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 修改后的crawl_mainpage方法 def crawl_mainpage(self, response): driver = response.request.meta['driver'] # 点击公司简介tab button = driver.find_element_by_xpath("//span[@title='Company Profile']") button.click() # 显式等待成立年份元素加载完成,最多等待10秒 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//td[contains(text(), 'Year Established')]/following-sibling::td/div/div/div")) ) except: pass # 从driver获取最新的页面源码,构造新的选择器 new_selector = Selector(text=driver.page_source) # 提取完成后再关闭driver driver.quit() # 用新的选择器提取数据,用extract_first()避免返回列表写入csv格式异常 yield { 'name': new_selector.xpath("//h1[@class='module-pdp-title']/text()").extract_first(), 'Year of Establishment': new_selector.xpath("//td[contains(text(), 'Year Established')]/following-sibling::td/div/div/div/text()").extract_first() }
额外注意事项
- 请确认
settings.py中已经正确开启scrapy-selenium中间件:
DOWNLOADER_MIDDLEWARES = { 'scrapy_selenium.SeleniumMiddleware': 800 }
- 如果频繁请求返回空,可以在
SeleniumRequest中增加wait_time参数,或者给driver配置随机UA、代理降低被反爬拦截的概率
内容的提问来源于stack exchange,提问作者TheGoldBerg
相关产品推荐
相关产品推荐

