Python HTTP Web代理服务器URL解析异常导致getaddrinfo错误及请求冻结问题解决
Fixing getaddrinfo Error & Freeze When Handling favicon.ico in Your Python Proxy Server
Let's break down why your proxy is freezing on favicon.ico requests and fix it step by step.
The Root Cause
When your browser loads http://127.0.0.1:8000/httpforever.com, it automatically sends a follow-up request for /favicon.ico to the proxy (since it thinks the proxy is the origin server). Your current URL parsing logic fails here:
- The request line for the icon is
GET /favicon.ico HTTP/1.1 filename = message.split()[1].partition("/")[2]gives you"favicon.ico"hostn = filename.replace("www.", "", 1).partition("/")[0]then setshostnto"favicon.ico"—which isn't a valid domain, hence thegetaddrinfoerror.- Worse, your code doesn't properly clean up the socket after this failure, leading to a freeze.
Fixed Implementation
Here's a revised version of your proxy with key fixes and improvements:
# Proxy Server from socket import * import sys import os # Create a server socket, bind it to a port and start listening tcpSerSock = socket(AF_INET, SOCK_STREAM) serverName = '0.0.0.0' serverPort = 8000 tcpSerSock.bind((serverName, serverPort)) tcpSerSock.listen(1) # Create cache directory to avoid clutter os.makedirs("cache", exist_ok=True) def parse_request(message): """Parse GET request to extract target host and resource path""" try: request_line = message.splitlines()[0] method, path, _ = request_line.split() if method != "GET": return None, None, "Only GET requests are supported" # Expect paths like /httpforever.com or /httpforever.com/favicon.ico if not path.startswith("/"): return None, None, "Invalid request path format" # Split into host and resource parts path_parts = path.lstrip("/").split("/", 1) hostn = path_parts[0] resource_path = "/" if len(path_parts) == 1 else "/" + path_parts[1] # Basic domain validation if "." not in hostn: return None, None, "Missing valid target host in request path" return hostn, resource_path, None except Exception as e: return None, None, f"Request parsing failed: {str(e)}" while 1: print('Ready to serve...') tcpCliSock, addr = tcpSerSock.accept() print('Received connection from:', addr) try: # Read more data and handle binary resources gracefully message = tcpCliSock.recv(4096).decode(errors="ignore") if not message: tcpCliSock.close() continue hostn, resource_path, error_msg = parse_request(message) if error_msg: print(f"Error: {error_msg}") # Send proper 400 response to client tcpCliSock.send(b"HTTP/1.0 400 Bad Request\r\n") tcpCliSock.send(b"Content-Type: text/html\r\n\r\n") tcpCliSock.send(f"<html><body><h1>400 Bad Request</h1><p>{error_msg}</p></body></html>".encode()) tcpCliSock.close() continue # Generate safe cache filename (replace slashes to avoid directory issues) cache_filename = f"cache/{hostn}_{resource_path.replace('/', '_')}" file_exists = False # Check cache first try: with open(cache_filename, "rb") as f: cached_data = f.read() file_exists = True tcpCliSock.send(cached_data) print(f"Served from cache: {cache_filename}") except IOError: # Cache miss: fetch from remote server remote_sock = None try: remote_sock = socket(AF_INET, SOCK_STREAM) remote_sock.connect((hostn, 80)) print(f"Connected to {hostn} successfully") # Build valid GET request with Host header (required by most servers) request_str = f"GET {resource_path} HTTP/1.0\r\nHost: {hostn}\r\n\r\n" remote_sock.sendall(request_str.encode()) # Read full response from remote server response_data = b"" while True: chunk = remote_sock.recv(4096) if not chunk: break response_data += chunk # Send response to client tcpCliSock.send(response_data) print(f"Fetched from {hostn}: {resource_path}") # Save to cache with open(cache_filename, "wb") as cache_file: cache_file.write(response_data) print(f"Saved to cache: {cache_filename}") except Exception as e: print(f"Failed to fetch resource: {str(e)}") # Send 502 Bad Gateway response tcpCliSock.send(b"HTTP/1.0 502 Bad Gateway\r\n") tcpCliSock.send(b"Content-Type: text/html\r\n\r\n") tcpCliSock.send(f"<html><body><h1>502 Bad Gateway</h1><p>{str(e)}</p></body></html>".encode()) finally: # Ensure remote socket is closed even if error occurs if remote_sock: remote_sock.close() tcpCliSock.close() except Exception as e: print(f"Unexpected error handling client {addr}: {str(e)}") if 'tcpCliSock' in locals() and not tcpCliSock._closed: tcpCliSock.close() tcpSerSock.close() sys.exit()
Key Improvements
Robust Request Parsing:
- The
parse_requestfunction explicitly extracts the target host and resource path from requests like/httpforever.comor/httpforever.com/favicon.ico. - Validates that the host is a proper domain (contains a dot) to avoid invalid
getaddrinfocalls.
- The
Proper Error Handling:
- Sends clear HTTP error responses (400, 502) to the client instead of silently failing.
- Ensures all sockets are closed even when errors occur, preventing freezes from unclosed connections.
Cache Management:
- Uses a dedicated
cachedirectory to store cached resources. - Generates safe filenames by replacing slashes, avoiding directory conflicts.
- Uses a dedicated
Better Resource Handling:
- Uses a larger buffer (4096 bytes) for more efficient data transfer.
- Handles binary resources (like images) gracefully with
decode(errors="ignore"). - Includes the required
Hostheader in requests to remote servers (many servers reject requests without this).
How to Test
- Access the main page via
http://127.0.0.1:8000/httpforever.com - Access the favicon directly via
http://127.0.0.1:8000/httpforever.com/favicon.ico
The proxy will now handle both requests without freezing or throwing getaddrinfo errors.
内容的提问来源于stack exchange,提问作者tzhou022
相关产品推荐
相关产品推荐

