如何使用Scrapy(Python2)爬取加载缓慢的动态页面?
老哥,你这个问题我太熟了!核心原因是Scrapy本身不执行JavaScript——它爬的是服务器直接返回的初始HTML,也就是那个带加载GIF的页面,而后续的内容是浏览器跑JS之后才生成的。你加的time.sleep()和download_delay根本没用,因为它们只是在请求前后等一等,完全触发不了页面的JS渲染流程。
给你两个靠谱的解决方案,按需选:
方案1:直接抓AJAX接口(优先选,效率高)
先打开浏览器的开发者工具(F12),切换到Network面板,刷新页面看看有没有XHR/Fetch请求——这些就是页面加载内容的接口。直接用Scrapy请求这些接口,比模拟浏览器快多了。举个Python2的例子:
import scrapy class QuotesSpider(scrapy.Spider): name = "quotes" def start_requests(self): # 把这里换成你找到的真实AJAX接口URL ajax_api_url = "https://example.com/api/load-quotes" yield scrapy.Request(url=ajax_api_url, callback=self.parse_ajax) def parse_ajax(self, response): # 接口一般返回JSON,直接解析就行 data = response.json() for quote_item in data.get('quotes', []): yield { 'text': quote_item.get('content'), 'author': quote_item.get('author_name') }
方案2:用Scrapy+Selenium模拟浏览器(复杂页面必备)
如果页面的动态逻辑特别绕,找不到AJAX接口,那就只能用真实浏览器来渲染页面了。Selenium可以帮你调用Chrome/Firefox,等页面加载完再抓内容。
首先得装依赖(Python2要装兼容版本的Selenium):
pip install selenium==2.53.6
然后修改你的爬虫代码:
import scrapy from selenium import webdriver from scrapy.http import HtmlResponse from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By class QuotesSpider(scrapy.Spider): name = "quotes" def __init__(self): # 这里要填你ChromeDriver的路径,记得下对应浏览器版本的驱动 self.driver = webdriver.Chrome(executable_path='/your/path/to/chromedriver') def start_requests(self): urls = ['https://example.com'] for url in urls: yield scrapy.Request(url=url, callback=self.parse) def parse(self, response): self.driver.get(response.url) # 别用time.sleep!用显式等待更靠谱——等目标元素出现再继续 try: # 最多等10秒,直到.quote这个元素加载出来(换成你要抓的元素选择器) WebDriverWait(self.driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, '.quote')) ) except Exception as e: self.logger.error(f"等待页面加载超时: {e}") return # 拿到渲染后的页面源码,转成Scrapy的Response对象,方便后续解析 rendered_html = self.driver.page_source new_response = HtmlResponse(url=response.url, body=rendered_html, encoding='utf-8') # 现在就可以正常用CSS/XPath提取内容了 for quote in new_response.css('.quote'): yield { 'text': quote.css('.text::text').extract_first(), 'author': quote.css('.author::text').extract_first() } def closed(self, reason): # 爬虫结束记得关浏览器,不然会留进程 self.driver.quit()
小提示:
- 尽量用显式等待替代
time.sleep(),前者是等元素出现再继续,后者是硬等,容易浪费时间或者错过加载时机。 - 如果可以的话,建议升级到Python3——Python2早就停更了,现在Scrapy的新版本、更方便的工具(比如Playwright)都只支持Python3,用起来舒服多了。
内容的提问来源于stack exchange,提问作者Codieroot
相关产品推荐
相关产品推荐

