Python Requests带Cookie的GET请求返回JSON数据缺失问题
Fixing Missing JSON Fields in Requests vs. Browser
Hey there, let's tackle this frustrating issue—when your code fetches JSON with fewer fields than the browser, it almost always comes down to differences in how the two are making requests. Let's break down the likely culprits and fix your code step by step.
Key Issues in Your Current Code
First, let's spot the small things that might be causing the problem:
- Manual cookie handling: You're extracting cookies from the session and passing them back explicitly, but the
requests.Sessionobject already manages cookies automatically. This can lead to conflicts or missing session data. - Request header mismatches: Browsers send a lot more headers than your code (like
User-Agent,Accept-Language, etc.), and many servers tailor responses based on these headers—often returning stripped-down data for non-browser clients. - Redundant JSON serialization: You're calling
json.dumps(secret)twice, and manually setting theContent-Typeheader when requests can handle this automatically.
Revised Code with Fixes
Here's an updated version of your code that addresses these issues:
import requests import json with requests.Session() as s: # Mimic browser headers to ensure full response data s.headers.update({ 'Accept': 'application/json, text/javascript, */*; q=0.01', 'Content-Type': 'application/json', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9' }) # Login step - use requests' built-in json parameter secret = {"client_id": "xxx", "client_secret": "yyy"} login_url = 'https://XXXX.yyy/sign_in' response = s.post(login_url, json=secret, verify=False) # Verify login succeeded (check status code or response content) if response.status_code != 200: print(f"Login failed with status: {response.status_code}") exit(1) # Fetch data - let session handle cookies automatically data_url = 'https://give.me/tests' print(f"Connecting to: {data_url}") response = s.get(data_url, verify=False) # Parse and save JSON properly try: json_data = response.json() print("Full JSON data retrieved:") print(json.dumps(json_data, indent=4)) with open('data.txt', 'w') as f: json.dump(json_data, f, indent=4) except json.JSONDecodeError: print("Error: Response is not valid JSON") print("Response content:", response.text)
Why This Works
- Automatic Cookie Management: The
Sessionobject keeps track of cookies set during login and automatically sends them with subsequent requests—no need to extract and reattach them manually. - Browser-like Headers: Adding a real
User-Agentand matchingAcceptheaders tells the server you're a legitimate browser client, so it returns the full JSON structure instead of a stripped-down version. - Simplified JSON Handling: Using
json=secretinstead ofdata=json.dumps(secret)lets requests handle serialization and set the correctContent-Typeheader automatically, reducing errors.
Additional Troubleshooting Tips
- Compare Request Headers: Open your browser's DevTools (Network tab), make the GET request, and copy all the headers sent by the browser. Add any missing ones to your
s.headers.update()call—sometimes headers likeRefererorOriginare required for full responses. - Check Response Status: Ensure your GET request returns a
200 OKstatus code. If it's a403or401, your session might not be properly authenticated (double-check the login response for errors). - Inspect Raw Responses: Print
response.textin your code to see exactly what's being returned—sometimes servers return HTML error pages instead of JSON when something's wrong with the request.
内容的提问来源于stack exchange,提问作者user3375601
相关产品推荐
相关产品推荐

