如何在不使用urllib系列及request库的情况下用Python下载文件
Solution: File Download with Python's Built-in
socket & ssl Modules Got it, since you can’t rely on high-level HTTP libraries like requests or urllib, we can drop down to Python’s core socket module—the low-level networking tool that powers all those libraries under the hood—to handle file downloads directly. I’ll cover both HTTP and HTTPS cases below, using only standard library modules you should have access to.
1. Basic HTTP File Download
This works for unencrypted HTTP endpoints. We’ll manually craft an HTTP GET request, parse the server’s response, and write the file content to disk.
import socket def download_http_file(url, save_path): # Parse the URL to extract host and file path if not url.startswith("http://"): url = f"http://{url}" host = url.split("://")[1].split("/")[0] path = "/" + "/".join(url.split("://")[1].split("/")[1:]) or "/" # Create socket and set timeout to avoid hanging sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM) sock.settimeout(10) try: sock.connect((host, 80)) # Send properly formatted HTTP GET request request = f"GET {path} HTTP/1.1\r\nHost: {host}\r\nConnection: close\r\n\r\n" sock.send(request.encode("utf-8")) # Receive response in chunks (avoids loading large files into memory) response = b"" while True: chunk = sock.recv(4096) if not chunk: break response += chunk # Split response headers from the actual file content header_end = response.find(b"\r\n\r\n") headers = response[:header_end].decode("utf-8") file_content = response[header_end+4:] # Validate we got a successful response if "HTTP/1.1 200 OK" not in headers: print(f"Error: Server returned non-success response\n{headers}") return # Write content to disk with open(save_path, "wb") as f: f.write(file_content) print(f"File saved successfully to {save_path}") except Exception as e: print(f"Download failed: {str(e)}") finally: sock.close() # Example usage download_http_file("example.com/sample.txt", "downloaded_sample.txt")
2. HTTPS File Download
For encrypted HTTPS connections, we just wrap the socket with Python’s ssl module (also part of the standard library) to handle TLS encryption.
import socket import ssl def download_https_file(url, save_path): # Parse URL for host and path if not url.startswith("https://"): url = f"https://{url}" host = url.split("://")[1].split("/")[0] path = "/" + "/".join(url.split("://")[1].split("/")[1:]) or "/" sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM) sock.settimeout(10) # Wrap socket with SSL/TLS to handle encrypted traffic ssl_sock = ssl.wrap_socket(sock) try: ssl_sock.connect((host, 443)) request = f"GET {path} HTTP/1.1\r\nHost: {host}\r\nConnection: close\r\n\r\n" ssl_sock.send(request.encode("utf-8")) # Receive and process response response = b"" while True: chunk = ssl_sock.recv(4096) if not chunk: break response += chunk header_end = response.find(b"\r\n\r\n") headers = response[:header_end].decode("utf-8") file_content = response[header_end+4:] if "HTTP/1.1 200 OK" not in headers: print(f"Error: Server returned non-success response\n{headers}") return with open(save_path, "wb") as f: f.write(file_content) print(f"Secure file saved successfully to {save_path}") except Exception as e: print(f"HTTPS download failed: {str(e)}") finally: ssl_sock.close() # Example usage download_https_file("example.com/secure_sample.txt", "secure_downloaded.txt")
Key Tips for Production Use
- Chunked Transfer Encoding: The above code works for most cases, but some servers use chunked encoding for large files. To handle this, you’ll need to parse the
Transfer-Encoding: chunkedheader, decode chunk sizes from hex, and read bytes accordingly. - Large Files: For extra-large files, avoid storing the entire content in memory—write each chunk to the file immediately after receiving it, instead of building a single
responsebyte string. - Header Validation: Add checks for
Content-Lengthif you want to verify the file size matches what the server claims.
内容的提问来源于stack exchange,提问作者Dark_John
相关产品推荐
相关产品推荐

