如何将BeautifulSoup获取的window._sharedData导入json.loads解析?
Got it, let's break this down for you. The data you've pulled from the site isn't valid JSON on its own—it's a JavaScript variable assignment statement. To use it with Python's json.loads(), you first need to strip out the non-JSON parts, then parse the remaining valid JSON object.
Here's a step-by-step solution:
Step 1: Clean the raw data
Your raw string starts with window._sharedData = , which isn't part of the JSON structure. We'll remove this prefix (and any trailing semicolon if present) to get a valid JSON string.
Step 2: Parse with json.loads()
Once you have the cleaned JSON string, you can use Python's json module to parse it into a dictionary, then extract whatever fields you need.
Full Code Example
import json # Your raw fetched data raw_data = '''window._sharedData = { "config": { "csrf_token": "DMjhhPBY0i6ZyMKYQPjMjxJhRD0gkRVQ", "viewer": null, "viewerId": null }, "country_code": "IN", "language_code": "en", "locale": "en_US" }''' # Step 1: Clean the data # Remove the "window._sharedData = " prefix cleaned_json = raw_data.replace('window._sharedData = ', '').strip() # Optional: Remove trailing semicolon if your raw data has one (e.g., ends with " };") if cleaned_json.endswith(';'): cleaned_json = cleaned_json[:-1].strip() # Step 2: Parse the JSON data = json.loads(cleaned_json) # Now extract your desired information csrf_token = data['config']['csrf_token'] country_code = data['country_code'] language_code = data['language_code'] print(f"CSRF Token: {csrf_token}") print(f"Country Code: {country_code}")
Explanation
- The
replace()method strips out the JavaScript variable declaration, leaving only the JSON object. - We add a check for a trailing semicolon because sometimes these JS assignments end with
;which would breakjson.loads(). - Once parsed, you can access values using standard dictionary key lookups, just like any other Python dict.
内容的提问来源于stack exchange,提问作者Harsha M V
相关产品推荐
相关产品推荐

