Scrapy使用FormRequest.from_response无法登录网页求助
It sounds like your login flow is getting stuck because Scrapy isn’t correctly interacting with the login form—let’s break down the most likely fixes to get you past that loop:
1. Target the Correct Login Form Explicitly
FormRequest.from_response defaults to selecting the first form on the page. If your login page has multiple forms (like login + forgot password), it might be submitting the wrong one. Fix this by explicitly targeting the login form using formxpath or formcss:
def parse(self, response): # Use the form's unique ID or attributes to select it return scrapy.FormRequest.from_response( response, formxpath='//form[@id="login-form"]', # Replace with your login form's actual XPath formdata={ 'username': 'your_username', # Match the input's "name" attribute from the form 'password': 'your_password' # Same here—double-check the input's name }, callback=self.after_login )
2. Verify Form Field Names Match Exactly
Double-check the name attributes of your username and password inputs in the login form. For example, some sites use email instead of username, or passwd instead of password. If these don’t match what you’re passing in formdata, the server will reject your login and redirect you back.
3. Don’t Overlook Hidden Fields (Like CSRF Tokens)
Most login forms include hidden CSRF tokens or session IDs that are required for submission. While FormRequest.from_response should auto-include these, if it’s not working, extract them manually:
def parse(self, response): # Extract the CSRF token using the input's name attribute csrf_token = response.css('input[name="csrfmiddlewaretoken"]::attr(value)').get() return scrapy.FormRequest( url='https://xyz/login', # Explicitly use the login URL formdata={ 'username': 'your_username', 'password': 'your_password', 'csrfmiddlewaretoken': csrf_token # Add the extracted token }, callback=self.after_login )
4. Add an after_login Callback to Validate Success
Create a separate callback to confirm if your login actually worked—this helps you debug whether the issue is with form submission or post-login navigation:
def after_login(self, response): # Check for a unique string that confirms successful login (e.g., "Welcome back") if "Welcome back" in response.text: self.logger.info("Login succeeded! Starting scrape.") # Now navigate to your target page yield scrapy.Request(url='https://xyz/target-page', callback=self.parse_target) else: self.logger.error("Login failed. Check credentials or form setup.") # Save the response to debug what's happening with open('login_failure.html', 'wb') as f: f.write(response.body)
5. Check User-Agent and Cookie Settings
Some servers block Scrapy’s default user-agent. Add a realistic one in your settings.py:
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
Also, ensure COOKIES_ENABLED = True (default in Scrapy) so cookies are persisted across requests—manual cookie injection often causes mismatched session data and loops.
By following these steps, you should be able to resolve the redirect loop and get your login working as expected.
内容的提问来源于stack exchange,提问作者karthik loganathan

