使用Python Socket发送HTTPS请求仅获部分内容的问题排查
Hey there! Let's figure out why you're only getting partial HTML content (even after logging in with cookies) when using raw socket and ssl modules for HTTPS requests. This is a common pitfall with low-level TCP/HTTP handling, so let's break down the fixes and checks step by step:
1. Fix Your HTTP Request Format
Your request is missing the critical header termination marker required by HTTP: two consecutive \r\n (carriage return + line feed) characters. Without this, the server might not recognize that your request is complete, leading to incomplete or delayed responses.
Update your request string to include proper line breaks and the required termination:
MySock.send(f"""GET / HTTP/1.1 Host: {HOST} Connection: close """.encode())
- We added
Connection: closeto tell the server to shut down the connection after sending the full response—this makes it easier to know when we've received all data. - The blank line at the end translates to
\r\n\r\n, which signals the end of the request headers.
2. Loop to Receive Full TCP Stream Data
TCP is a stream-based protocol, meaning a single recv() call only grabs whatever data is currently in the socket buffer—it doesn't guarantee you'll get the entire response in one go. You need to keep receiving chunks until the server closes the connection (returns an empty byte string).
Replace your single recv() call with a loop to collect all data:
full_response = b"" while True: chunk = MySock.recv(4096) # 4096 is a standard buffer size for such operations if not chunk: break # Server closed the connection, no more data full_response += chunk # Now write the complete response to file with open('PythonOrg.html', 'w', encoding='utf-8') as file: file.write(full_response.decode())
This will stitch together all the chunks of data sent by the server, giving you the full HTML content instead of just a partial snippet.
3. Handle Chunked Transfer Encoding (If Needed)
Many modern websites use Transfer-Encoding: chunked instead of a Content-Length header. This means the response is split into multiple "chunks" with size markers. While the loop above will still work (thanks to Connection: close), if you want to properly parse chunked responses without relying on the connection closing, you'd need to:
- Parse the response headers to check if
Transfer-Encoding: chunkedis present - For each chunk, read the size line, then read that many bytes of data
- Stop when you hit a chunk with size 0
But for most cases, the Connection: close + loop approach is simpler and works perfectly.
4. Update Deprecated SSL Code
Your current ssl.wrap_socket() call is deprecated. For better security and compatibility, use ssl.create_default_context() instead. Here's how to refactor that part of your code:
import socket import ssl HOST = 'www.python.org' PORT = 443 # Create a secure SSL context context = ssl.create_default_context() # Use context managers to auto-close sockets/SSL connections with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as sock: with context.wrap_socket(sock, server_hostname=HOST) as my_sock: my_sock.connect((HOST, PORT)) # Send your fixed request here my_sock.send(f"""GET / HTTP/1.1 Host: {HOST} Connection: close """.encode()) # Receive full response with the loop from step 2 full_response = b"" while True: chunk = my_sock.recv(4096) if not chunk: break full_response += chunk with open('PythonOrg.html', 'w', encoding='utf-8') as file: file.write(full_response.decode())
Using with statements ensures sockets are properly closed, preventing resource leaks.
内容的提问来源于stack exchange,提问作者rSim

