You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy简单验证码求解:寻求测试页面的可行实现方案

Scrapy 处理验证码的入门方案及实操示例

Hey there, I’ve looked into your problem and put together a practical, beginner-friendly Scrapy solution to handle the captcha on your test login page. Let’s break this down step by step:

1. Install Required Dependencies

First, you’ll need to set up the tools we’ll use for image processing and OCR (Optical Character Recognition):

  • Install Scrapy if you haven’t already: pip install scrapy
  • For image handling: pip install pillow
  • For OCR: pip install pytesseract
  • Important: You also need the Tesseract engine itself. On Ubuntu, run sudo apt install tesseract-ocr; on macOS, use brew install tesseract; on Windows, download the installer from the official site and add it to your system PATH.

2. Core Workflow

The basic idea for handling captchas in Scrapy is straightforward:

  • Fetch the login page to get the captcha image.
  • Process the image to clean it up (improves OCR accuracy).
  • Use OCR to read the captcha text.
  • Submit the login form with the recognized captcha.

3. Full Spider Code Example

Here’s a complete Scrapy spider tailored to your test page. I’ve added comments to explain each step:

import scrapy
import pytesseract
from PIL import Image
from io import BytesIO

class CaptchaLoginSpider(scrapy.Spider):
    name = 'captcha_login'
    start_urls = ['http://145.100.108.148/login3/']

    def parse(self, response):
        # Grab the captcha image URL from the login page
        captcha_img_relative_url = response.css('img[src*="captcha"]::attr(src)').get()
        # Convert relative URL to absolute (in case it's not full path)
        captcha_img_url = response.urljoin(captcha_img_relative_url)

        # Send a request to fetch the captcha image, passing along the login page response
        yield scrapy.Request(
            captcha_img_url,
            callback=self.parse_captcha,
            meta={'login_response': response}
        )

    def parse_captcha(self, response):
        # Step 1: Process the image to boost OCR accuracy
        img = Image.open(BytesIO(response.body))
        # Convert to grayscale
        img = img.convert('L')
        # Apply binary thresholding (adjust 127 if needed for your captcha)
        threshold = 127
        img = img.point(lambda x: 0 if x < threshold else 255, '1')

        # Step 2: Use Tesseract to read the captcha text
        captcha_text = pytesseract.image_to_string(img).strip()
        self.logger.info(f"Recognized captcha text: {captcha_text}")

        # Step 3: Prepare the login form data
        # Replace with your actual test username/password
        form_payload = {
            'username': 'test_user',
            'password': 'test_pass',
            'captcha': captcha_text
        }

        # Submit the login form using the original login page response
        yield scrapy.FormRequest.from_response(
            response.meta['login_response'],
            formdata=form_payload,
            callback=self.verify_login
        )

    def verify_login(self, response):
        # Check if login was successful (adjust the check based on your page's content)
        if 'Welcome' in response.text or '登录成功' in response.text:
            self.logger.info("✅ Login succeeded!")
            # Add code here to scrape post-login content
        else:
            self.logger.warning("❌ Login failed—likely bad captcha recognition. Retrying...")
            # Retry by fetching the login page again
            yield scrapy.Request(self.start_urls[0], callback=self.parse)

4. Key Configuration & Tips

  • Tesseract Path (Windows Users): If you’re on Windows, add this line at the top of your spider to point to the Tesseract executable:
    pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'
    
  • Improve OCR Accuracy: If the captcha has noise (like lines/dots), tweak the threshold value or add more image processing steps (e.g., using PIL’s filter() method for noise reduction).
  • Avoid Getting Blocked: Add a delay between requests in your settings.py:
    DOWNLOAD_DELAY = 2
    
  • Retry Logic: The example includes a basic retry on login failure, but you can also use Scrapy’s built-in RetryMiddleware for more robust handling.

5. If Local OCR Isn’t Cutting It

If Tesseract struggles with more complex captchas (twisted text, heavy noise), you have options:

  • Train a custom Tesseract model using samples of your target captcha.
  • Use a dedicated captcha-breaking library (some require training your own model).
  • Use a third-party captcha-solving API (paid services work well for tricky cases).

内容的提问来源于stack exchange,提问作者Kevin C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:26:48