使用Python Requests调用API时出现无效URL问题求助
Fixing "Invalid URL" Error When Fetching Scroll-Loaded Content API with Requests
Let's break down why you're getting that "Invalid URL" response and fix the script step by step:
Key Issues Identified
- Incorrect Host Header: The
Hostheader should only contain the domain name (ie.api.atom.nowtv.com), not the fullhttps://www.nowtv.comURL. Servers use this header to route requests correctly, so the wrong value here will immediately break your request. - Unescaped HTML Entities in URL: When you copied the URL from Chrome DevTools, the
&characters were escaped as&(this is how HTML displays&to avoid parsing issues). The API server needs the raw&to split query parameters, so leaving these escaped will make the URL unrecognizable.
Corrected Script
import requests import cloudscraper # Fix: Replace & with raw & in the URL url = 'https://ie.api.atom.nowtv.com/adapter-atlas/v3/query/node?slug=/entertainment/collections/all-entertainment&represent=(items[take=60](items(items[select_list=iceberg])))' session = requests.session() # Fix: Set Host to the correct API domain (no https:// prefix) session.headers = { 'Host': 'ie.api.atom.nowtv.com', 'Connection': 'keep-alive', 'Accept': 'application/json, text/javascript, */*', 'X-Requested-With': 'XMLHttpRequest', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/72.0.3626.119 Safari/537.36', 'Referer': 'https://www.nowtv.com', 'Accept-Encoding': 'gzip, deflate', 'Accept-Language': 'en-GB,en-US;q=0.9,en;q=0.8' } # Initialize cloudscraper with your session scraper = cloudscraper.create_scraper(sess=session) r = scraper.get(url) # Try parsing as JSON first (since the API should return JSON) try: data = r.json() print(data) except ValueError: # Fallback to printing raw text if JSON parsing fails print(r.text) session.close()
Additional Notes
- When copying URLs or headers from DevTools, always double-check for HTML-escaped characters like
&,<, or>— these need to be converted back to their original form for requests to work. - Using
r.json()instead ofr.contentwill automatically parse the response into a Python dictionary, making it much easier to work with the program data you're trying to fetch. - If you still run into issues, verify that all headers match exactly what Chrome sends (you can use the "Copy as cURL" feature in DevTools and convert it to requests syntax if needed).
内容的提问来源于stack exchange,提问作者gdogg371
相关产品推荐
相关产品推荐

