You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决服务器拦截网页爬取引发的HTTP 403及代码执行错误

解决网页爬取403拦截与inscriptis适配问题

咱们一步步拆解你遇到的问题,同时给出可直接运行的解决方案:

一、先解决urllib的headers参数错误问题

你之前直接给urlopen()传headers参数会报错,是因为urlopen()本身不支持这个参数——得先通过Request对象来封装请求头,再把这个对象传给urlopen()。这样既解决了TypeError,又能通过请求头模拟浏览器,避免403拦截。

示例代码:

from urllib.request import Request, urlopen
from inscriptis import get_text

target_url = "https://economictimes.indiatimes.com"
# 核心是添加User-Agent,模拟真实浏览器请求
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8"
}

# 先构建Request对象,把headers传进去
req = Request(target_url, headers=request_headers)
# 发起请求并读取HTML内容
with urlopen(req) as response:
    html_content = response.read().decode('utf-8')

# 用inscriptis提取文本
extracted_text = get_text(html_content)
print(extracted_text)

二、用requests模块的正确打开方式

你之前遇到的AttributeError,是因为直接把requests的Response对象传给了get_text()——这个函数需要的是HTML字符串或者类文件对象,不是Response实例。只需要取出Response里的文本内容再传入就可以了:

示例代码:

import requests
from inscriptis import get_text

target_url = "https://economictimes.indiatimes.com"
request_headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 发起带headers的请求
response = requests.get(target_url, headers=request_headers)
# 先检查请求是否成功,避免后续处理失败
response.raise_for_status()

# 取出HTML文本传给get_text()
extracted_text = get_text(response.text)
print(extracted_text)

三、进阶:应对更严格的反爬策略

如果加了User-Agent还是遇到403,可以试试这些技巧:

  • 补充更多浏览器请求头,比如Accept-Language、Referer,让请求更接近真实用户;
  • 添加请求间隔,用time.sleep(1)之类的代码,避免短时间内频繁请求触发拦截;
  • 若网站有Cookie验证,可以先用浏览器访问一次,把Cookie复制到请求头里;
  • 必要时可以使用代理IP,但这需要额外的代理资源支持。

内容的提问来源于stack exchange,提问作者Kristada673

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:18:47