You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy Playwright报错ValueError: Page.evaluate事件循环不匹配

问题:Scrapy+Playwright调用AntiCaptcha后执行页面操作报错

我在通过AntiCaptcha API获取验证码解决方案后,尝试用Playwright点击提交按钮,运行代码时触发如下错误:

line 514, in wrap_api_call
    raise rewrite_error(error, f"{parsed_st['apiName']}: {error}") from None
ValueError: Page.evaluate: The future belongs to a different loop than the one specified as the loop argument

我的完整代码如下:

import scrapy
from anticaptchaofficial.recaptchav2proxyless import recaptchaV2Proxyless
from scrapy_playwright.page import PageMethod

class RecaptchaSpider(scrapy.Spider):
    name = "recaptcha_spider"
    start_urls = ["https://www.google.com/recaptcha/api2/demo"]
    delay_after_submit = 5  # Time in seconds to wait after submit

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                dont_filter=True,
                meta={
                    "playwright": True,
                    "playwright_include_page": True,
                },
                callback=self.solve_captcha,
            )

    async def solve_captcha(self, response):
        # Step 1: Solve the CAPTCHA using AntiCaptcha API
        api_key = "my api key"  # Replace with your actual API key
        solver = recaptchaV2Proxyless()
        solver.set_key(api_key)
        solver.set_website_url(response.url)
        solver.set_website_key("6Le-wvkSAAAAAPBMRTvw0Q4Muexq9bi0DJwx_mJ-")  # Replace with Google reCAPTCHA site key

        captcha_solution = solver.solve_and_return_solution()
        if captcha_solution:
            self.logger.info(f"Solved CAPTCHA: {captcha_solution}")
            
            # Step 2: Execute JavaScript to fill the CAPTCHA solution and click the submit button
            page = response.meta["playwright_page"]
            await page.evaluate(f'document.getElementById("g-recaptcha-response").innerHTML = "{captcha_solution}";')  # Replace with the correct ID if different
            await page.click("button[type='submit']")  # Replace with the correct selector for the submit button
            
            # Wait for the next page or any other action you want to perform
            await page.wait_for_timeout(self.delay_after_submit)
            
            # Step 3: Handle the next step after clicking submit (if necessary)
            # You might want to retrieve the new page content here
            new_content = await page.content()
            self.logger.info("New page content after submit:")
            self.logger.info(new_content)
        else:
            self.logger.error("Failed to solve CAPTCHA.")

报错原因

recaptchaV2Proxyless的solve_and_return_solution()是同步阻塞方法,调用它会直接阻塞Scrapy的异步事件循环,导致后续Playwright的页面操作(如page.evaluate、page.click)处于不同的事件循环上下文,从而触发该错误。

修复方案

将同步的验证码求解操作放到线程池中执行,避免阻塞异步循环,修改后的代码如下:

import scrapy
from anticaptchaofficial.recaptchav2proxyless import recaptchaV2Proxyless
from scrapy_playwright.page import PageMethod
from twisted.internet.threads import deferToThread

class RecaptchaSpider(scrapy.Spider):
    name = "recaptcha_spider"
    start_urls = ["https://www.google.com/recaptcha/api2/demo"]
    delay_after_submit = 5  # 提交后的等待时间(秒)

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                dont_filter=True,
                meta={
                    "playwright": True,
                    "playwright_include_page": True,
                },
                callback=self.solve_captcha,
            )

    async def solve_captcha(self, response):
        # Step 1: 在线程池中执行同步的验证码求解
        api_key = "my api key"  # 替换为你的AntiCaptcha API密钥
        solver = recaptchaV2Proxyless()
        solver.set_key(api_key)
        solver.set_website_url(response.url)
        solver.set_website_key("6Le-wvkSAAAAAPBMRTvw0Q4Muexq9bi0DJwx_mJ-")

        # 用deferToThread把同步方法包装成异步可等待对象
        captcha_solution = await deferToThread(solver.solve_and_return_solution)
        
        if captcha_solution:
            self.logger.info(f"验证码已解决: {captcha_solution}")
            
            # Step 2: 填充验证码并提交
            page = response.meta["playwright_page"]
            # 改用参数传递方式注入token,避免JS注入风险
            await page.evaluate(
                '(token) => document.getElementById("g-recaptcha-response").innerHTML = token',
                captcha_solution
            )
            await page.click("button[type='submit']")
            
            # 等待页面加载或跳转
            await page.wait_for_timeout(self.delay_after_submit)
            
            # Step 3: 获取提交后的页面内容
            new_content = await page.content()
            self.logger.info("提交后的页面内容:")
            self.logger.info(new_content)
        else:
            self.logger.error("验证码求解失败")

额外优化提示

  • 避免直接将验证码token拼接进JavaScript字符串,改用page.evaluate的参数传递方式,防止潜在的JavaScript注入风险。
  • 如果AntiCaptcha提供官方异步SDK,优先使用异步版本,性能会比线程池方案更优。

内容的提问来源于stack exchange,提问作者boyenec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 11:34:51