You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy使用FormRequest.from_response无法登录网页求助

Troubleshooting Your Scrapy Login Loop Issue

It sounds like your login flow is getting stuck because Scrapy isn’t correctly interacting with the login form—let’s break down the most likely fixes to get you past that loop:

1. Target the Correct Login Form Explicitly

FormRequest.from_response defaults to selecting the first form on the page. If your login page has multiple forms (like login + forgot password), it might be submitting the wrong one. Fix this by explicitly targeting the login form using formxpath or formcss:

def parse(self, response):
    # Use the form's unique ID or attributes to select it
    return scrapy.FormRequest.from_response(
        response,
        formxpath='//form[@id="login-form"]',  # Replace with your login form's actual XPath
        formdata={
            'username': 'your_username',  # Match the input's "name" attribute from the form
            'password': 'your_password'   # Same here—double-check the input's name
        },
        callback=self.after_login
    )

2. Verify Form Field Names Match Exactly

Double-check the name attributes of your username and password inputs in the login form. For example, some sites use email instead of username, or passwd instead of password. If these don’t match what you’re passing in formdata, the server will reject your login and redirect you back.

3. Don’t Overlook Hidden Fields (Like CSRF Tokens)

Most login forms include hidden CSRF tokens or session IDs that are required for submission. While FormRequest.from_response should auto-include these, if it’s not working, extract them manually:

def parse(self, response):
    # Extract the CSRF token using the input's name attribute
    csrf_token = response.css('input[name="csrfmiddlewaretoken"]::attr(value)').get()
    return scrapy.FormRequest(
        url='https://xyz/login',  # Explicitly use the login URL
        formdata={
            'username': 'your_username',
            'password': 'your_password',
            'csrfmiddlewaretoken': csrf_token  # Add the extracted token
        },
        callback=self.after_login
    )

4. Add an after_login Callback to Validate Success

Create a separate callback to confirm if your login actually worked—this helps you debug whether the issue is with form submission or post-login navigation:

def after_login(self, response):
    # Check for a unique string that confirms successful login (e.g., "Welcome back")
    if "Welcome back" in response.text:
        self.logger.info("Login succeeded! Starting scrape.")
        # Now navigate to your target page
        yield scrapy.Request(url='https://xyz/target-page', callback=self.parse_target)
    else:
        self.logger.error("Login failed. Check credentials or form setup.")
        # Save the response to debug what's happening
        with open('login_failure.html', 'wb') as f:
            f.write(response.body)

Some servers block Scrapy’s default user-agent. Add a realistic one in your settings.py:

USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'

Also, ensure COOKIES_ENABLED = True (default in Scrapy) so cookies are persisted across requests—manual cookie injection often causes mismatched session data and loops.

By following these steps, you should be able to resolve the redirect loop and get your login working as expected.

内容的提问来源于stack exchange,提问作者karthik loganathan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:02:28