You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Buildozer环境下,Python不导入模块如何下载HTML数据?

Download HTML from a URL Without Requests/Urllib (Using Python's Built-in Socket)

Hey there! Let me break this down for you:

Strictly speaking, there's no way to make network requests in Python without importing any modules at all—network operations aren't part of Python's core syntax, so you need to rely on standard library tools. But the good news is you can use Python's built-in socket module, which is part of the core standard library. Buildozer should support this out of the box, no C recipes required!

How it works (step-by-step):

  • Manually parse the URL to split out the host and path
  • Create a socket connection to the target host on port 80 (the default HTTP port)
  • Send a raw HTTP GET request
  • Receive the full response, then split off the HTTP headers to extract the raw HTML content

Full Code Example:

def download_html(url):
    # Handle URLs missing the http:// prefix
    if not url.startswith('http://'):
        url = 'http://' + url
    
    # Split URL into host domain and request path
    host_end_idx = url.find('/', 7)
    if host_end_idx == -1:
        host = url[7:]
        path = '/'
    else:
        host = url[7:host_end_idx]
        path = url[host_end_idx:]
    
    # Import Python's built-in socket module (no extra dependencies)
    import socket
    # Initialize a TCP socket
    sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
    
    try:
        # Connect to the target host on port 80
        sock.connect((host, 80))
        
        # Send a raw HTTP GET request
        request = f"GET {path} HTTP/1.1\r\nHost: {host}\r\nConnection: close\r\n\r\n"
        sock.send(request.encode('utf-8'))
        
        # Receive the full response in chunks
        response = b''
        while True:
            chunk = sock.recv(1024)
            if not chunk:
                break
            response += chunk
    finally:
        # Ensure the socket is closed regardless of errors
        sock.close()
    
    # Separate HTTP headers from the actual HTML content
    header_separator = b'\r\n\r\n'
    sep_idx = response.find(header_separator)
    if sep_idx != -1:
        # Decode HTML content (handle encoding errors gracefully)
        html_content = response[sep_idx + len(header_separator):].decode('utf-8', errors='replace')
        return html_content
    else:
        # Fallback if header separator isn't found
        return response.decode('utf-8', errors='replace')

# Usage example
if __name__ == "__main__":
    page_html = download_html("example.com")
    print(page_html)

Important Notes:

  • This implementation only supports HTTP URLs (not HTTPS). If you need HTTPS support, you can extend it with Python's built-in ssl module (also part of the standard library) to wrap the socket—just let me know if you need that!
  • It doesn't handle redirects (3xx status codes) or error responses (4xx/5xx) by default. You can add logic to check the HTTP status code in the response headers if you need to handle these cases.
  • The errors='replace' flag prevents crashes if the target site uses a non-UTF-8 encoding. You can adjust this to match the specific encoding of the site you're targeting.
  • Since socket is part of Python's core standard library, Buildozer will include it automatically when packaging your app—no extra configuration or C recipes needed.

内容的提问来源于stack exchange,提问作者TellMeWhy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:20:10