urllib.request.urlopen无法打开部分HTTPS URL,求HTML提取替代方法
问题解决:urllib无法访问otcmarkets.com的方案及替代HTML提取方法
一、urllib无法访问的原因及修复方法
otcmarkets.com这类网站会检测请求的User-Agent字段,urllib默认的请求头会暴露这是Python程序发起的请求,从而被反爬机制拦截。修复方法是手动添加模拟浏览器的请求头:
import urllib.request # 构造请求头,模拟Chrome浏览器 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } req = urllib.request.Request(url="https://www.otcmarkets.com", headers=headers) html = urllib.request.urlopen(req) # 读取响应内容 content = html.read().decode('utf-8') print(content[:500]) # 打印前500个字符验证
如果还是无法访问,可以尝试添加更多请求头字段,进一步模拟真实浏览器请求:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' }
二、更简便的HTML提取替代方法
1. 使用requests库
requests库比urllib更简洁易用,同样需要设置请求头绕过反爬:
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } response = requests.get("https://www.otcmarkets.com", headers=headers) response.encoding = 'utf-8' # 指定编码 html_content = response.text print(html_content[:500])
2. 使用httpx库(支持同步/异步)
httpx是现代的HTTP客户端,API和requests类似,支持HTTP/2,同样可以设置请求头:
import httpx headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } with httpx.Client() as client: response = client.get("https://www.otcmarkets.com", headers=headers) html_content = response.text print(html_content[:500])
注:如果网站有更严格的反爬机制(比如验证码、Cookie验证),可能需要用selenium或playwright模拟浏览器行为,但这类方法更重,适合复杂场景。
内容的提问来源于stack exchange,提问作者Gal Eindor
相关产品推荐
相关产品推荐

