如何让Anaconda Python 3.6的urllib调用curl执行HTTP请求?
实现urllib调用curl执行请求的方案
核心思路:自定义urllib Handler替换底层请求逻辑
通过继承urllib.request.BaseHandler并实现http_open/https_open方法,将urllib的请求转发给curl执行,再把curl的响应包装成urllib兼容的响应对象,现有代码无需修改即可直接使用。
自定义CurlHTTPHandler实现
import urllib.request import subprocess import io import re class CurlHTTPHandler(urllib.request.BaseHandler): def __init__(self, curl_path=r"C:\Program Files\Git\mingw64\bin\curl.exe"): self.curl_path = curl_path def _execute_curl(self, req): # 构造curl基础命令 cmd = [ self.curl_path, "--ntlm", "-i", # 捕获响应头 "-X", req.get_method(), ] # 处理代理(从请求头中提取,或根据现有代码的代理设置调整) proxy = req.get_header('Proxy') if proxy: cmd.extend(["--proxy", proxy, "--proxy-ntlm"]) # 添加所有请求头(排除已单独处理的Proxy头) for header_name, header_value in req.headers.items(): if header_name.lower() != 'proxy': cmd.extend(["-H", f"{header_name}: {header_value}"]) # 处理请求体(POST/PUT等方法) if req.data is not None: cmd.extend(["--data-binary", "@-"]) # 从标准输入读取请求体 # 添加目标URL cmd.append(req.full_url) # 执行curl命令 process = subprocess.Popen( cmd, stdin=subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=False # 用二进制模式处理,兼容非文本响应 ) # 传入请求体(如果存在) stdout, stderr = process.communicate(input=req.data) if process.returncode != 0: raise urllib.request.URLError(f"Curl执行错误: {stderr.decode('utf-8')}") # 分离响应头和响应体 header_sep = b'\r\n\r\n' if b'\r\n\r\n' in stdout else b'\n\n' header_end = stdout.find(header_sep) headers_bytes = stdout[:header_end] body_bytes = stdout[header_end + len(header_sep):] # 解析状态码 status_line = headers_bytes.split(b'\n')[0].decode('utf-8') status_code = int(re.search(r'HTTP/\d+\.\d+ (\d+)', status_line).group(1)) # 解析响应头 headers = urllib.request.parse_headers(headers_bytes.decode('utf-8')) # 包装响应体为类文件对象 body_file = io.BytesIO(body_bytes) # 返回urllib兼容的响应对象 return urllib.response.addinfourl( body_file, headers, req.get_full_url(), status_code ) def http_open(self, req): return self._execute_curl(req) def https_open(self, req): return self._execute_curl(req)
启用自定义Handler
将自定义Handler注册到urllib的全局Opener中,后续所有urllib请求都会自动使用curl执行:
# 初始化自定义Handler(根据git-bash实际路径调整curl_path) curl_handler = CurlHTTPHandler(curl_path=r"C:\Program Files\Git\usr\bin\curl.exe") # 构建并安装全局Opener opener = urllib.request.build_opener(curl_handler) urllib.request.install_opener(opener) # 现有代码无需修改,直接使用urllib即可 response = urllib.request.urlopen("http://internal-api.example.com/endpoint") print(response.read().decode('utf-8'))
关键细节说明
- curl路径:务必根据git-bash的实际安装路径调整
curl_path,常见路径还有C:\Program Files\Git\usr\bin\curl.exe - 代理处理:如果现有代码通过
ProxyHandler设置代理,需确保请求头中包含Proxy字段,或在Handler中直接传入固定代理地址 - 认证信息:若需要指定NTLM用户名密码,可在curl命令中添加
--user username:password参数,比如在cmd列表中加入["--user", "your-username:your-password"] - 二进制响应:使用
text=False模式处理curl输出,确保图片、压缩文件等二进制内容正常返回 - 错误处理:可根据需求扩展异常捕获逻辑,处理curl返回的不同错误码
无需额外安装的替代方案
如果不想依赖curl,可尝试通过ctypes调用Windows系统的winhttp.dll实现NTLM认证,但该方案需要编写大量底层调用代码,复杂度较高。核心思路如下:
- 使用
ctypes加载winhttp.dll - 调用
WinHttpOpen、WinHttpConnect等函数建立连接 - 设置NTLM认证选项并发送请求
- 解析响应并包装成urllib兼容对象
该方案无需额外软件,但实现成本远高于curl封装,仅在无法使用curl时考虑。
内容的提问来源于stack exchange,提问作者briconaut
相关产品推荐
相关产品推荐

