使用Scrapy Playwright爬取Ajax页面时无输出的问题求助
Scrapy Playwright爬取Ajax页面时无输出的问题求助
我最近尝试用Scrapy Playwright爬取https://www.scrapethissite.com/pages/ajax-javascript/这个网站的内容,但遇到了问题——执行爬虫后没有任何输出,想请大家帮忙排查一下问题所在。
我准备爬取的页面HTML代码截图如下:
以下是我编写的爬虫代码:
import scrapy from scrapy_playwright.page import PageMethod class OscarSpider(scrapy.Spider): name = "OscarSpider" def start_requests(self): yield scrapy.Request( url="https://www.scrapethissite.com/pages/ajax-javascript/", callback=self.parse, meta={ "playwright": True, "playwright_include_page": True, "playwright_page_methods": [ PageMethod("wait_for_selector", "a#2010"), # 确保按钮存在 PageMethod("click", "a#2010"), # 点击按钮 PageMethod("wait_for_selector", "tr.film"), # 等待数据加载 PageMethod("evaluate", "window.scrollTo(0, document.body.scrollHeight)"), PageMethod("wait_for_timeout", 6000) # 等待AJAX数据加载 ] } ) async def parse(self, response): for row in response.css("tr.film"): yield { "title": row.css("td.film-title::text").get(default="").strip(), "nominations": row.css("td.film-nominations::text").get(default="").strip(), "awards": row.css("td.film-awards::text").get(default="").strip(), }
当我执行以下命令运行爬虫时,没有得到任何输出:
scrapy crawl OscarSpider -O Oscar.json
我期望得到的JSON输出示例如下(展示部分数据):
Title Nominations Awards The King's Speech 12 4 Inception 8 4
麻烦大家帮忙看看问题出在哪里,谢谢!
备注:内容来源于stack exchange,提问作者Nitish K
相关产品推荐
相关产品推荐

