解决服务器拦截网页爬取引发的HTTP 403及代码执行错误
解决网页爬取403拦截与inscriptis适配问题
咱们一步步拆解你遇到的问题,同时给出可直接运行的解决方案:
一、先解决urllib的headers参数错误问题
你之前直接给urlopen()传headers参数会报错,是因为urlopen()本身不支持这个参数——得先通过Request对象来封装请求头,再把这个对象传给urlopen()。这样既解决了TypeError,又能通过请求头模拟浏览器,避免403拦截。
示例代码:
from urllib.request import Request, urlopen from inscriptis import get_text target_url = "https://economictimes.indiatimes.com" # 核心是添加User-Agent,模拟真实浏览器请求 request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8" } # 先构建Request对象,把headers传进去 req = Request(target_url, headers=request_headers) # 发起请求并读取HTML内容 with urlopen(req) as response: html_content = response.read().decode('utf-8') # 用inscriptis提取文本 extracted_text = get_text(html_content) print(extracted_text)
二、用requests模块的正确打开方式
你之前遇到的AttributeError,是因为直接把requests的Response对象传给了get_text()——这个函数需要的是HTML字符串或者类文件对象,不是Response实例。只需要取出Response里的文本内容再传入就可以了:
示例代码:
import requests from inscriptis import get_text target_url = "https://economictimes.indiatimes.com" request_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 发起带headers的请求 response = requests.get(target_url, headers=request_headers) # 先检查请求是否成功,避免后续处理失败 response.raise_for_status() # 取出HTML文本传给get_text() extracted_text = get_text(response.text) print(extracted_text)
三、进阶:应对更严格的反爬策略
如果加了User-Agent还是遇到403,可以试试这些技巧:
- 补充更多浏览器请求头,比如
Accept-Language、Referer,让请求更接近真实用户; - 添加请求间隔,用
time.sleep(1)之类的代码,避免短时间内频繁请求触发拦截; - 若网站有Cookie验证,可以先用浏览器访问一次,把Cookie复制到请求头里;
- 必要时可以使用代理IP,但这需要额外的代理资源支持。
内容的提问来源于stack exchange,提问作者Kristada673
相关产品推荐
相关产品推荐

