Scrapy Playwright报错ValueError: Page.evaluate事件循环不匹配
问题:Scrapy+Playwright调用AntiCaptcha后执行页面操作报错
我在通过AntiCaptcha API获取验证码解决方案后,尝试用Playwright点击提交按钮,运行代码时触发如下错误:
line 514, in wrap_api_call raise rewrite_error(error, f"{parsed_st['apiName']}: {error}") from None ValueError: Page.evaluate: The future belongs to a different loop than the one specified as the loop argument
我的完整代码如下:
import scrapy from anticaptchaofficial.recaptchav2proxyless import recaptchaV2Proxyless from scrapy_playwright.page import PageMethod class RecaptchaSpider(scrapy.Spider): name = "recaptcha_spider" start_urls = ["https://www.google.com/recaptcha/api2/demo"] delay_after_submit = 5 # Time in seconds to wait after submit def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, dont_filter=True, meta={ "playwright": True, "playwright_include_page": True, }, callback=self.solve_captcha, ) async def solve_captcha(self, response): # Step 1: Solve the CAPTCHA using AntiCaptcha API api_key = "my api key" # Replace with your actual API key solver = recaptchaV2Proxyless() solver.set_key(api_key) solver.set_website_url(response.url) solver.set_website_key("6Le-wvkSAAAAAPBMRTvw0Q4Muexq9bi0DJwx_mJ-") # Replace with Google reCAPTCHA site key captcha_solution = solver.solve_and_return_solution() if captcha_solution: self.logger.info(f"Solved CAPTCHA: {captcha_solution}") # Step 2: Execute JavaScript to fill the CAPTCHA solution and click the submit button page = response.meta["playwright_page"] await page.evaluate(f'document.getElementById("g-recaptcha-response").innerHTML = "{captcha_solution}";') # Replace with the correct ID if different await page.click("button[type='submit']") # Replace with the correct selector for the submit button # Wait for the next page or any other action you want to perform await page.wait_for_timeout(self.delay_after_submit) # Step 3: Handle the next step after clicking submit (if necessary) # You might want to retrieve the new page content here new_content = await page.content() self.logger.info("New page content after submit:") self.logger.info(new_content) else: self.logger.error("Failed to solve CAPTCHA.")
报错原因
recaptchaV2Proxyless的solve_and_return_solution()是同步阻塞方法,调用它会直接阻塞Scrapy的异步事件循环,导致后续Playwright的页面操作(如page.evaluate、page.click)处于不同的事件循环上下文,从而触发该错误。
修复方案
将同步的验证码求解操作放到线程池中执行,避免阻塞异步循环,修改后的代码如下:
import scrapy from anticaptchaofficial.recaptchav2proxyless import recaptchaV2Proxyless from scrapy_playwright.page import PageMethod from twisted.internet.threads import deferToThread class RecaptchaSpider(scrapy.Spider): name = "recaptcha_spider" start_urls = ["https://www.google.com/recaptcha/api2/demo"] delay_after_submit = 5 # 提交后的等待时间(秒) def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, dont_filter=True, meta={ "playwright": True, "playwright_include_page": True, }, callback=self.solve_captcha, ) async def solve_captcha(self, response): # Step 1: 在线程池中执行同步的验证码求解 api_key = "my api key" # 替换为你的AntiCaptcha API密钥 solver = recaptchaV2Proxyless() solver.set_key(api_key) solver.set_website_url(response.url) solver.set_website_key("6Le-wvkSAAAAAPBMRTvw0Q4Muexq9bi0DJwx_mJ-") # 用deferToThread把同步方法包装成异步可等待对象 captcha_solution = await deferToThread(solver.solve_and_return_solution) if captcha_solution: self.logger.info(f"验证码已解决: {captcha_solution}") # Step 2: 填充验证码并提交 page = response.meta["playwright_page"] # 改用参数传递方式注入token,避免JS注入风险 await page.evaluate( '(token) => document.getElementById("g-recaptcha-response").innerHTML = token', captcha_solution ) await page.click("button[type='submit']") # 等待页面加载或跳转 await page.wait_for_timeout(self.delay_after_submit) # Step 3: 获取提交后的页面内容 new_content = await page.content() self.logger.info("提交后的页面内容:") self.logger.info(new_content) else: self.logger.error("验证码求解失败")
额外优化提示
- 避免直接将验证码token拼接进JavaScript字符串,改用
page.evaluate的参数传递方式,防止潜在的JavaScript注入风险。 - 如果AntiCaptcha提供官方异步SDK,优先使用异步版本,性能会比线程池方案更优。
内容的提问来源于stack exchange,提问作者boyenec
相关产品推荐
相关产品推荐

