You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python带缓存HTTP Web代理服务器的请求解析问题及getaddrinfo错误解决求助

看起来你遇到的问题核心是请求路径和主机名的解析逻辑有漏洞,尤其是当浏览器自动请求favicon.ico这类资源时,代码要么构造了错误的请求路径,要么解析出无效的主机名,导致getaddrinfo失败,进而让socket连接卡住。下面我会一步步拆解问题并给出修复方案:

1. 错误的请求路径构造

当前代码中,无论请求的资源是什么,你都构造了固定的GET / HTTP/1.0\r\n\r\n请求,这只在请求根路径时有效。当请求favicon.ico或其他子路径时,目标服务器根本不知道你要请求什么资源,自然无法返回正确内容,甚至可能导致连接异常。

修复方案:
从filename中拆分出主机名和资源路径:

  • 主机名是filename中第一个/之前的部分
  • 资源路径是filename中第一个/之后的部分(如果没有则为/)

同时添加Host请求头——很多服务器需要这个字段来正确处理请求:

# 拆分主机名和资源路径
if "/" in filename:
    hostn, path = filename.split("/", 1)
    path = f"/{path}"
else:
    hostn = filename
    path = "/"

# 构造正确的GET请求
request_str = f"GET {path} HTTP/1.0\r\nHost: {hostn}\r\n\r\n"

2. 无效主机名的处理

当浏览器直接请求http://127.0.0.1:8000/favicon.ico(而非带主机名的完整路径)时,filename会变成favicon.ico,此时hostn就是favicon.ico——这显然不是一个有效的域名,getaddrinfo会直接失败,导致程序卡住。

修复方案:
捕获getaddrinfo错误并返回明确的400错误,同时提示无效主机名:

# 记得导入gaierror
from socket import gaierror

# ... 其他代码 ...
try:
    c.connect((hostn, 80))
    # ... 后续请求逻辑
except gaierror:
    error_msg = b"HTTP/1.0 400 Bad Request\r\n\r\nInvalid hostname provided"
    tcpCliSock.send(error_msg)
    print(f"Failed to resolve hostname: {hostn}")
except Exception as E:
    print("Illegal request")
    print(E)

3. 缓存文件的正确保存与读取

你的代码注释掉了缓存保存逻辑,而且即使启用,之前的处理也不够严谨:

  • 应该把服务器返回的**完整响应(包括HTTP头)**保存到缓存,这样缓存命中时可以直接发送完整响应,避免手动构造错误的Content-Type
  • 文件名要替换所有/,避免创建子目录

修复方案:
取消注释缓存保存代码,用二进制模式写入:

# 保存到缓存
cache_filename = filename.replace("/", "_")
with open(cache_filename, "wb") as tempFile:
    tempFile.write(buf_data)

缓存命中时直接读取二进制文件并发送:

try:
    with open(cache_filename, "rb") as f:
        outputdata = f.read()
        fileExists = True
        # 直接发送完整响应(包含正确的Content-Type)
        tcpCliSock.send(outputdata)
        print("Read from cache")
except IOError:
    # ... 缓存未命中的逻辑

4. 字符串发送的编码问题

之前的代码中,从缓存读取文本文件后直接发送字符串会报错(socket.send()只接受bytes类型),用上面的二进制读取方式可以自动解决这个问题;如果仍需文本读取,要手动编码:

tcpCliSock.send(line.encode())

修改后的完整代码

from socket import *
import sys

# Create a server socket, bind it to a port and start listening
tcpSerSock = socket(AF_INET, SOCK_STREAM)
serverName = '0.0.0.0'
serverPort = 8000
tcpSerSock.bind((serverName, serverPort))
tcpSerSock.listen(1)

while 1:
    # Start receiving data from the client
    print('Ready to serve...')
    tcpCliSock, addr = tcpSerSock.accept()
    print('Received a connection from:', addr)
    
    try:
        message = tcpCliSock.recv(1024).decode()
        if not message:
            tcpCliSock.close()
            continue
        print(message)
        
        # Extract the filename from the given message
        request_path = message.split()[1]
        filename = request_path[1:] if request_path.startswith('/') else request_path
        print(f"filename: {filename}")
        
        # 生成缓存文件名
        cache_filename = filename.replace("/", "_")
        fileExists = False
        
        try:
            # Check whether the file exist in the cache
            with open(cache_filename, "rb") as f:
                outputdata = f.read()
                fileExists = True
                # ProxyServer finds a cache hit and sends the full response
                tcpCliSock.send(outputdata)
                print("Read from cache")
        except IOError:
            if not fileExists:
                # Create a socket on the proxyserver
                c = socket(AF_INET, SOCK_STREAM)
                # 拆分主机名和资源路径
                if "/" in filename:
                    hostn, path = filename.split("/", 1)
                    path = f"/{path}"
                else:
                    hostn = filename
                    path = "/"
                print(f"hostn:{hostn}, path:{path}")
                
                try:
                    c.connect((hostn, 80))
                    print("Successfully connected")
                    # 构造包含Host头的正确请求
                    request_str = f"GET {path} HTTP/1.0\r\nHost: {hostn}\r\n\r\n"
                    c.sendall(request_str.encode())
                    
                    buf_data = b""
                    # Read the response into buffer
                    while True:
                        data = c.recv(1024)
                        if not data:
                            break
                        buf_data += data
                    
                    # Send response to client
                    tcpCliSock.send(buf_data)
                    
                    # Save to cache
                    with open(cache_filename, "wb") as tempFile:
                        tempFile.write(buf_data)
                    print(f"Saved to cache: {cache_filename}")
                    
                except gaierror:
                    error_msg = b"HTTP/1.0 400 Bad Request\r\n\r\nInvalid hostname provided"
                    tcpCliSock.send(error_msg)
                    print(f"Failed to resolve hostname: {hostn}")
                except Exception as E:
                    print("Illegal request")
                    print(E)
                finally:
                    c.close()
            else:
                nf_header = b"HTTP/1.0 404 Not Found\r\n\r\n<h1>404 Not Found</h1>"
                tcpCliSock.send(nf_header)
    except Exception as e:
        print(f"Error handling request: {e}")
    finally:
        tcpCliSock.close()

tcpSerSock.close()
sys.exit()

测试验证

现在用你的测试链接http://127.0.0.1:8000/httpforever.com访问,浏览器会自动请求http://127.0.0.1:8000/httpforever.com/favicon.ico,此时代码会正确解析主机名httpforever.com和路径/favicon.ico,向目标服务器请求资源并缓存,不会再出现冻结或getaddrinfo错误。

内容的提问来源于stack exchange,提问作者tzhou022

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 09:47:27