使用Scrapy带凭据爬取Glassdoor:FORMDATA缺失问题求助
Hey Arnaud, let's work through the issues you're hitting with your Scrapy login script for Glassdoor, plus help you track down that missing form data you need.
First: Fix the Critical Errors in Your Current Code
Your script has a few straightforward mistakes that are likely causing problems right off the bat:
- Incorrect
allowed_domains: This should only be the root domain, not a full URL. Change it toallowed_domains = ['glassdoor.co.uk']— your current value will make Scrapy block valid requests to Glassdoor. - Broken
start_urls: You've got a duplicatehttp://prefix (http://https://...). Fix it tostart_urls = ['https://www.glassdoor.co.uk/profile/login_input.htm']. - Missing hidden form fields: Glassdoor's login form almost certainly uses dynamic hidden fields (like CSRF tokens or session IDs) that your script isn't capturing. Just passing username/password won't work unless you include these.
Finding That Missing Form Data (The 302/POST Request Issue)
You mentioned not seeing the Form Data section for the 302 request — here's why and how to fix it:
- 302 is the redirect response, not the login request: The 302 is what Glassdoor sends back after processing your login, not the POST request you sent to log in. To see your actual login submission, enable the Preserve log checkbox in Chrome DevTools' Network tab before you log in. This keeps all requests in the log, even after redirects, so you can find the original POST request (it'll likely have a 200 status code, not 302).
- Check for XHR/AJAX requests: Glassdoor might handle login via an AJAX call instead of a traditional form submit. Switch to the XHR tab in DevTools when logging in — look for requests to paths like
/api/loginor similar. The "Payload" tab of that request will show you all the form data you need (including hidden fields).
Revised Scrapy Script Example
Here's an updated version of your script that addresses the above issues, including extracting hidden form fields:
import scrapy from scrapy.http import FormRequest from scrapy.utils.response import open_in_browser class GdSpider(scrapy.Spider): name = 'gd' allowed_domains = ['glassdoor.co.uk'] start_urls = ['https://www.glassdoor.co.uk/profile/login_input.htm'] def parse(self, response): # Extract hidden CSRF token (adjust the selector to match Glassdoor's actual form) csrf_token = response.css('input[name="csrfToken"]::attr(value)').get() if not csrf_token: self.logger.warning("Failed to find CSRF token - check the form selector!") # Use FormRequest.from_response to auto-populate hidden fields, then override credentials return FormRequest.from_response( response, formxpath='//form[contains(@action, "login")]', # Target the correct login form formdata={ 'username': 'your_username', 'password': 'your_password', 'csrfToken': csrf_token # Include the hidden token if required }, callback=self.scrape_pages, dont_filter=True ) def scrape_pages(self, response): # Verify login success by checking for a unique element (like "Dashboard" in the page) if "Dashboard" in response.text: self.logger.info("Login successful!") open_in_browser(response) else: self.logger.warning("Login failed - double-check credentials or form fields!")
Extra Tips for Glassdoor Scraping
Glassdoor has strict anti-bot measures, so keep these in mind:
- Set a reasonable
DOWNLOAD_DELAYin yoursettings.py(e.g.,DOWNLOAD_DELAY = 3) to avoid triggering rate limits. - Spoof a real user-agent string to mimic a browser — add this to
settings.py:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' - If you hit CAPTCHAs or blocks, you might need to use rotating proxies or a CAPTCHA-solving service (though this adds complexity).
内容的提问来源于stack exchange,提问作者Arnaud C
相关产品推荐
相关产品推荐

