You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python HTTP Web代理服务器URL解析异常导致getaddrinfo错误及请求冻结问题解决

Fixing getaddrinfo Error & Freeze When Handling favicon.ico in Your Python Proxy Server

Let's break down why your proxy is freezing on favicon.ico requests and fix it step by step.

The Root Cause

When your browser loads http://127.0.0.1:8000/httpforever.com, it automatically sends a follow-up request for /favicon.ico to the proxy (since it thinks the proxy is the origin server). Your current URL parsing logic fails here:

  • The request line for the icon is GET /favicon.ico HTTP/1.1
  • filename = message.split()[1].partition("/")[2] gives you "favicon.ico"
  • hostn = filename.replace("www.", "", 1).partition("/")[0] then sets hostn to "favicon.ico"—which isn't a valid domain, hence the getaddrinfo error.
  • Worse, your code doesn't properly clean up the socket after this failure, leading to a freeze.

Fixed Implementation

Here's a revised version of your proxy with key fixes and improvements:

# Proxy Server
from socket import *
import sys
import os

# Create a server socket, bind it to a port and start listening
tcpSerSock = socket(AF_INET, SOCK_STREAM)
serverName = '0.0.0.0'
serverPort = 8000
tcpSerSock.bind((serverName, serverPort))
tcpSerSock.listen(1)

# Create cache directory to avoid clutter
os.makedirs("cache", exist_ok=True)

def parse_request(message):
    """Parse GET request to extract target host and resource path"""
    try:
        request_line = message.splitlines()[0]
        method, path, _ = request_line.split()
        
        if method != "GET":
            return None, None, "Only GET requests are supported"
        
        # Expect paths like /httpforever.com or /httpforever.com/favicon.ico
        if not path.startswith("/"):
            return None, None, "Invalid request path format"
        
        # Split into host and resource parts
        path_parts = path.lstrip("/").split("/", 1)
        hostn = path_parts[0]
        resource_path = "/" if len(path_parts) == 1 else "/" + path_parts[1]
        
        # Basic domain validation
        if "." not in hostn:
            return None, None, "Missing valid target host in request path"
        
        return hostn, resource_path, None
    except Exception as e:
        return None, None, f"Request parsing failed: {str(e)}"

while 1:
    print('Ready to serve...')
    tcpCliSock, addr = tcpSerSock.accept()
    print('Received connection from:', addr)
    
    try:
        # Read more data and handle binary resources gracefully
        message = tcpCliSock.recv(4096).decode(errors="ignore")
        if not message:
            tcpCliSock.close()
            continue
        
        hostn, resource_path, error_msg = parse_request(message)
        if error_msg:
            print(f"Error: {error_msg}")
            # Send proper 400 response to client
            tcpCliSock.send(b"HTTP/1.0 400 Bad Request\r\n")
            tcpCliSock.send(b"Content-Type: text/html\r\n\r\n")
            tcpCliSock.send(f"<html><body><h1>400 Bad Request</h1><p>{error_msg}</p></body></html>".encode())
            tcpCliSock.close()
            continue
        
        # Generate safe cache filename (replace slashes to avoid directory issues)
        cache_filename = f"cache/{hostn}_{resource_path.replace('/', '_')}"
        file_exists = False
        
        # Check cache first
        try:
            with open(cache_filename, "rb") as f:
                cached_data = f.read()
                file_exists = True
                tcpCliSock.send(cached_data)
                print(f"Served from cache: {cache_filename}")
        except IOError:
            # Cache miss: fetch from remote server
            remote_sock = None
            try:
                remote_sock = socket(AF_INET, SOCK_STREAM)
                remote_sock.connect((hostn, 80))
                print(f"Connected to {hostn} successfully")
                
                # Build valid GET request with Host header (required by most servers)
                request_str = f"GET {resource_path} HTTP/1.0\r\nHost: {hostn}\r\n\r\n"
                remote_sock.sendall(request_str.encode())
                
                # Read full response from remote server
                response_data = b""
                while True:
                    chunk = remote_sock.recv(4096)
                    if not chunk:
                        break
                    response_data += chunk
                
                # Send response to client
                tcpCliSock.send(response_data)
                print(f"Fetched from {hostn}: {resource_path}")
                
                # Save to cache
                with open(cache_filename, "wb") as cache_file:
                    cache_file.write(response_data)
                print(f"Saved to cache: {cache_filename}")
                
            except Exception as e:
                print(f"Failed to fetch resource: {str(e)}")
                # Send 502 Bad Gateway response
                tcpCliSock.send(b"HTTP/1.0 502 Bad Gateway\r\n")
                tcpCliSock.send(b"Content-Type: text/html\r\n\r\n")
                tcpCliSock.send(f"<html><body><h1>502 Bad Gateway</h1><p>{str(e)}</p></body></html>".encode())
            finally:
                # Ensure remote socket is closed even if error occurs
                if remote_sock:
                    remote_sock.close()
        
        tcpCliSock.close()
    except Exception as e:
        print(f"Unexpected error handling client {addr}: {str(e)}")
        if 'tcpCliSock' in locals() and not tcpCliSock._closed:
            tcpCliSock.close()

tcpSerSock.close()
sys.exit()

Key Improvements

  1. Robust Request Parsing:

    • The parse_request function explicitly extracts the target host and resource path from requests like /httpforever.com or /httpforever.com/favicon.ico.
    • Validates that the host is a proper domain (contains a dot) to avoid invalid getaddrinfo calls.
  2. Proper Error Handling:

    • Sends clear HTTP error responses (400, 502) to the client instead of silently failing.
    • Ensures all sockets are closed even when errors occur, preventing freezes from unclosed connections.
  3. Cache Management:

    • Uses a dedicated cache directory to store cached resources.
    • Generates safe filenames by replacing slashes, avoiding directory conflicts.
  4. Better Resource Handling:

    • Uses a larger buffer (4096 bytes) for more efficient data transfer.
    • Handles binary resources (like images) gracefully with decode(errors="ignore").
    • Includes the required Host header in requests to remote servers (many servers reject requests without this).

How to Test

  • Access the main page via http://127.0.0.1:8000/httpforever.com
  • Access the favicon directly via http://127.0.0.1:8000/httpforever.com/favicon.ico

The proxy will now handle both requests without freezing or throwing getaddrinfo errors.

内容的提问来源于stack exchange,提问作者tzhou022

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 06:40:20