You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

urllib库User Agent问题导致无法获取网页HTML的求助

解决urllib.error.HTTPError: 403 Forbidden(User Agent相关问题)

嘿,我太懂这种被403拒之门外的感觉了!这基本就是网站的反爬机制在起作用——它检测到你的请求不是来自常规浏览器,直接把你挡回去了。你已经意识到要加User Agent是对的,但代码里可能有几个小细节没处理好,我帮你一步步解决:

1. 补全你的User Agent字符串

你代码里的Firefox UA后面是...,这肯定不行,网站识别不了残缺的UA。得用一个完整、有效的浏览器User Agent,比如最新版Firefox或者Chrome的都可以,举两个例子:

  • Firefox: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/117.0
  • Chrome: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36

2. 修正Request和urlopen的调用方式

你导入了Request,但调用时的写法有点小问题,应该先构造完整的Request对象,再传给urlopen。我给你写个修正后的完整代码:

from urllib.request import urlopen, Request
from bs4 import BeautifulSoup as soup
import urllib.error

my_url = 'https://hsreplay.net/meta/#tab=matchups&sortBy=winrate'
# 用完整的User Agent
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/117.0'}

try:
    # 先构造带headers的Request
    req = Request(my_url, headers=headers)
    uClient = urlopen(req)
    page_html = uClient.read()
    uClient.close()
    
    # 解析HTML
    page_soup = soup(page_html, "html.parser")
    print("网页内容获取成功!")
    # 这里可以继续写你的解析逻辑
except urllib.error.HTTPError as e:
    print(f"HTTP错误: {e}")
except Exception as e:
    print(f"其他错误: {e}")

3. 进阶优化:模拟更真实的浏览器请求

如果加了UA还是不行,那可以多加点请求头,比如Accept、Accept-Language,让请求看起来更像真实浏览器:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/117.0',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5'
}

4. 备选方案:用requests库更省心

如果urllib用着麻烦,推荐试试requests库,它的语法更简洁,会话管理也更方便,代码大概是这样:

import requests
from bs4 import BeautifulSoup as soup

my_url = 'https://hsreplay.net/meta/#tab=matchups&sortBy=winrate'
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/117.0'}

try:
    response = requests.get(my_url, headers=headers)
    response.raise_for_status()  # 自动抛出HTTP错误
    page_soup = soup(response.text, "html.parser")
    print("成功获取内容!")
except requests.exceptions.HTTPError as e:
    print(f"HTTP错误: {e}")
except Exception as e:
    print(f"其他错误: {e}")

最后提醒一句:爬取前最好看看网站的robots.txt,确保你的行为符合网站的使用条款哦!

内容的提问来源于stack exchange,提问作者Javier Jiménez de la Jara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:57:11