Python实现URL提示文件保存循环请求(禁用Selenium)
Got it, let's work through this step by step with simple Python code—no Selenium needed. We'll fix the 403 error, loop through URLs to save files, and parse that weird callback-wrapped JSON from livePubJSON.js.
Step 1: Beat the HTTP 403 Forbidden Error
The 403 is almost certainly because the server is flagging your request as non-browser traffic. You already have the correct YouTube-derived User-Agent, so we'll center our headers around that. I'll add a couple more standard headers to make the request look more legitimate, but the User-Agent is the critical piece here.
Define your headers like this:
import requests import json # Your provided User-Agent plus extra standard browser-like headers headers = { 'User-Agent': 'Mozilla/5.0 (X11; Linux i868) AppleWebKit/537.17 (KHTML, like Gecko) Chrome/24.0.1312.27 Safari/537.17', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5' }
Step 2: Loop Through URLs & Save Files
Now we'll build a simple loop to hit each target URL, validate the response, and save the content to a file. Use binary mode (wb) for non-text files (like videos or images) or text mode (w) for files like JSON/JS.
Example code:
# List of URLs you need to process url_list = [ 'https://example.com/target-file-1', 'https://example.com/target-file-2', # Add all your URLs here ] for index, url in enumerate(url_list): try: # Send request with our browser-like headers response = requests.get(url, headers=headers) response.raise_for_status() # Trigger error if status code isn't 200 # Generate a simple filename (customize this based on URL or needs) filename = f"saved_item_{index+1}.txt" # Swap extension to .mp4/.json as needed # Save the content to disk with open(filename, 'wb') as f: f.write(response.content) print(f"✅ Saved {filename} successfully") except requests.exceptions.RequestException as e: print(f"❌ Failed to process {url}: {str(e)}")
Step 3: Parse the livePubJSON.js Callback-Wrapped JSON
The livePubJSON.js file has that annoying wrapper (like superagentCallback1518811920564( {"eu":{"...")) which breaks standard JSON parsing. We just need to strip off that callback function wrapper to get valid JSON.
Here's how to do it:
# Read the local livePubJSON.js file with open('livePubJSON.js', 'r') as f: js_content = f.read() # Strip the callback wrapper: find the first '(' and last ')' start_idx = js_content.find('(') + 1 end_idx = js_content.rfind(')') clean_json_str = js_content[start_idx:end_idx].strip() # Parse into a Python dictionary try: parsed_data = json.loads(clean_json_str) print("✅ JSON parsed successfully!") # Access your data, e.g., parsed_data['eu'] print("Sample data:", parsed_data['eu']) except json.JSONDecodeError as e: print(f"❌ Failed to parse JSON: {str(e)}")
Combine It All (Fetch & Parse in One Go)
If you need to fetch the livePubJSON.js directly via API and parse it without saving to disk first, combine the steps:
# URL of the livePubJSON.js file js_api_url = 'https://example.com/livePubJSON.js' # Fetch and parse in one step response = requests.get(js_api_url, headers=headers) response.raise_for_status() js_content = response.text start_idx = js_content.find('(') + 1 end_idx = js_content.rfind(')') clean_json_str = js_content[start_idx:end_idx].strip() parsed_data = json.loads(clean_json_str) print("Final parsed JSON data:", parsed_data)
Quick reminder: Always make sure you're respecting the website's robots.txt and terms of service when fetching or scraping content.
内容的提问来源于stack exchange,提问作者CAI

