BeautifulSoup爬虫无法写入网页内容,Waze实时地图RSS爬取异常
Hey there, let's break down your two crawler issues one by one and work through fixes for each:
Issue 1: Only writing "'b'" instead of JSON content from Waze RSS Feed
That "'b'" string is a dead giveaway—you're likely writing the string representation of a bytes object directly to file, instead of decoding it first. Here's how to fix this:
Key Fixes & Steps:
- Decode the response properly: When you fetch content with
requests,response.contentreturns raw bytes. To get readable text (or JSON), useresponse.text(already decoded to UTF-8 by default) or explicitly decode bytes with.decode('utf-8'). - Handle JSON correctly: If the RSS feed returns JSON data, use
response.json()to parse it into a Python dictionary, then convert it back to a formatted string for writing. - Validate your request: Add
response.raise_for_status()to catch silent HTTP errors (like 403 Forbidden or 500 Server Error) that might be returning empty/invalid content.
Example Code:
import requests import json waze_rss_url = "your_waze_rss_feed_url" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "application/json, application/rss+xml" # Match expected content type } try: response = requests.get(waze_rss_url, headers=headers) response.raise_for_status() # Crash early if request fails # Try parsing as JSON first json_data = response.json() with open("waze_data.json", "w", encoding="utf-8") as f: json.dump(json_data, f, indent=4, ensure_ascii=False) # Pretty-print for readability except ValueError: # Fallback if it's actual RSS XML instead of JSON from bs4 import BeautifulSoup soup = BeautifulSoup(response.text, "xml") with open("waze_rss.xml", "w", encoding="utf-8") as f: f.write(soup.prettify()) except Exception as e: print(f"Request failed: {str(e)}")
Issue 2: BeautifulSoup crawler can't write webpage content
If your BeautifulSoup script isn't writing content, it's usually one of three issues: failed requests, incorrect content extraction, or encoding problems. Here's how to diagnose and fix:
Key Fixes & Steps:
- Verify the request succeeded: Always check the HTTP status code and use
response.raise_for_status()to catch errors like blocked requests (common with missing user-agent headers). - Use the right parser: For HTML pages, use
"html.parser"or"lxml"(install withpip install lxmlfor better performance). For XML, use"xml". - Extract content correctly: Use
soup.get_text()for plain text, orsoup.prettify()to write formatted HTML. Avoid empty selectors (e.g.,soup.find("nonexistent-tag")would returnNone). - Specify encoding when writing: Always use
encoding="utf-8"in youropen()call to avoid garbled or missing content due to character set mismatches.
Example Code:
import requests from bs4 import BeautifulSoup target_url = "your_target_webpage_url" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: response = requests.get(target_url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") # Option 1: Write plain text content page_text = soup.get_text(strip=True, separator="\n") with open("webpage_text.txt", "w", encoding="utf-8") as f: f.write(page_text) # Option 2: Write formatted HTML # with open("webpage_html.html", "w", encoding="utf-8") as f: # f.write(soup.prettify()) print("Content written successfully!") except Exception as e: print(f"Error: {str(e)}")
Quick Troubleshooting Tips:
- Print
response.status_codeandresponse.text[:500](first 500 characters) to check if you're actually getting the content you expect. - If you're blocked, try rotating user-agent headers or adding a small delay between requests.
- For dynamic content (loaded via JavaScript), BeautifulSoup won't help—you'll need tools like Selenium or Playwright to render the page first.
内容的提问来源于stack exchange,提问作者Farnaz Khaghani
相关产品推荐
相关产品推荐

