使用Python requests下载网页获错误内容,求解决%20编码问题
Hey there! I’ve dealt with similar encoding headaches before, so let’s walk through how to fix this using just the requests library—no BeautifulSoup needed.
First: Avoid Manual %20 Encoding (Let Requests Handle It)
Your URL has %20 (the URL-encoded space), but sometimes manually encoding parameters can lead to unexpected issues. Instead, let requests handle parameter encoding automatically by passing them as a dictionary:
import requests # Base URL without the query string base_url = "https://some.com/ATS/cgi-bin/view.pl" # Define your parameters with actual spaces (requests will encode them correctly) query_params = { "axm_f_qtr_ins_chk_bd_sds status": "P01" } # Send the request with params instead of hardcoding %20 response = requests.get(base_url, params=query_params)
This way, you don’t have to worry about incorrect encoding—requests follows standard URL encoding rules for you.
Second: Fix Garbled Content by Setting the Right Encoding
Most of the time, garbled text happens because requests guesses the wrong character encoding for the response. Here’s how to fix it:
Check the response headers for charset:
The server usually tells you the encoding via theContent-Typeheader. You can access it and set the encoding explicitly:# Get the charset from headers (if available) content_type = response.headers.get('Content-Type', '') if 'charset=' in content_type: charset = content_type.split('charset=')[-1] response.encoding = charsetUse requests’ apparent encoding as a fallback:
If the headers don’t specify a charset,requestscan guess the actual encoding from the content itself withapparent_encoding:response.encoding = response.apparent_encoding # Now get the decoded text decoded_content = response.textManual decoding as a last resort:
If the above still doesn’t work, grab the raw byte content and try decoding with common encodings (like UTF-8, GBK, or ISO-8859-1):raw_content = response.content try: decoded_content = raw_content.decode('utf-8') except UnicodeDecodeError: try: decoded_content = raw_content.decode('gbk') # Common for Chinese content except UnicodeDecodeError: decoded_content = raw_content.decode('iso-8859-1') # Fallback encoding
Final Notes
Start with the parameter dictionary approach first—it eliminates any chance of manual encoding mistakes. Then tweak the encoding settings based on the response you get. This should resolve both the URL encoding concern and the garbled content issue.
内容的提问来源于stack exchange,提问作者user2381346

