Scrapy脚本登录Amscan网站遭遇HTTP 400错误求助
Hey, I see you're hitting a 400 error when trying to log into Amscan's B2B portal using Scrapy. Let's break down what's going wrong and fix it step by step.
First off, that 400 error is almost certainly because your request format doesn't match what the server expects. The login endpoint https://b2b.amscan.com/auth/login accepts JSON payloads, but you're using scrapy.FormRequest—which automatically encodes data as application/x-www-form-urlencoded (form data) instead of the required application/json. That's a classic mismatch that servers hate.
Here's how to fix this, plus some extra tweaks to make sure your request works:
1. Switch from FormRequest to Request for JSON Payloads
Instead of using FormRequest, use the base scrapy.Request class. This lets you send raw JSON as the request body, which is exactly what the endpoint wants.
2. Fix Your JSON Format
You had a string payload with single quotes earlier—JSON requires double quotes for all keys and string values. Also, make sure your credentials are correctly included (I noticed you had my_username/my_password in formdata but hardcoded them in the payload string; let's keep it consistent).
3. Handle 400 Responses for Debugging
By default, Scrapy ignores 400 responses. Adding meta={'handle_httpstatus_list': [400]} lets you inspect the response content even if it's a 400, which is super helpful for debugging.
Updated Working Code
import scrapy from scrapy.shell import inspect_response import json class Amscan(scrapy.Spider): name = 'amscan' start_urls = ['https://www.amscan.com/'] def parse(self, response): login_url = 'https://b2b.amscan.com/auth/login' # Properly formatted JSON payload (double quotes required!) payload = { 'login': { 'mode': 'email', 'username': 'info@soaringfinancial.com', 'password': 'PassSoaring$123' }, 'redirect': '#!/' } headers = { 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.88 Safari/537.36', 'content-type': 'application/json;charset=UTF-8', 'authority': 'b2b.amscan.com', 'sec-ch-ua': '"Google Chrome";v="87", " Not;A Brand";v="99", "Chromium";v="87"', 'accept': 'application/json, text/plain, */*', 'sec-ch-ua-mobile': '?0', 'sec-fetch-site': 'same-origin', 'sec-fetch-mode': 'cors', 'sec-fetch-dest': 'empty', 'accept-language': 'en-US,en;q=0.9', } # Send POST request with JSON body yield scrapy.Request( login_url, method='POST', body=json.dumps(payload), # Serialize payload to JSON string headers=headers, callback=self.login, meta={'handle_httpstatus_list': [400]} # Let's see what the server says even if 400 ) def login(self, response): # Print response content to debug any issues print("Response content:", response.text) inspect_response(response, self)
Extra Things to Check If It Still Fails
- Compare with Browser Request: Open Chrome DevTools (F12), go to the Network tab, and manually log into the site. Check the login request's Request Payload and Headers—make sure your Scrapy request matches exactly (sometimes sites add hidden fields or require specific headers you might have missed).
- Cookie Handling: Scrapy automatically manages cookies, but if the site requires a session cookie from the initial
amscan.comvisit, make sure your spider is following that correctly (the default start_urls should handle this, but you can verify in the response cookies). - Credentials: Double-check that your username and password are correct—sometimes 400 errors can mask authentication issues, so verifying with manual login first is a good step.
内容的提问来源于stack exchange,提问作者Viren Ramani

