You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python获取Zendesk托管URL内容时遇403 Forbidden错误的代码解决方案

解决Zendesk URL 403 Forbidden错误的代码层方案

原始代码与错误信息

Python代码

url = 'https://support.abc.com/hc...'
html_content = fetch_content(url)
if html_content:
    text_content = html_to_text(html_content.decode('utf-8'))
    # Output directory and file name
    output_pdf_file = os.path.join('output', 'output.pdf')
    if not os.path.exists('output'):
        os.makedirs('output')
        save_text_to_pdf(text_content, output_pdf_file)
        print(f'PDF saved to {output_pdf_file}')

错误信息

Error fetching https://support.abc.com/hc....: HTTP Error 403: Forbidden

代码层解决方法

  • 添加浏览器模拟User-Agent:Zendesk的反爬机制常拦截非浏览器标识的请求。修改fetch_content的请求头,加入符合浏览器特征的User-Agent,示例:

    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    # 假设fetch_content用requests实现,修改为:
    html_content = requests.get(url, headers=headers).content
    
  • 携带Cookie验证:部分Zendesk页面需要有效Cookie才能访问。手动在浏览器打开目标URL,复制浏览器DevTools中Cookie字段的内容,添加到请求头:

    headers = {
        'Cookie': 'your_cookie_string_from_browser',
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    

    注意Cookie存在过期可能,需要定期更新。

  • 添加Referer头:部分网站会校验请求来源,模拟从Zendesk主站跳转的请求:

    headers = {
        'Referer': 'https://support.abc.com/',
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    
  • 使用会话保持:用requests.Session()维持会话状态,模拟浏览器的持续连接,避免每次请求被判定为新请求:

    import requests
    
    session = requests.Session()
    session.headers.update({
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    })
    response = session.get(url)
    if response.status_code == 200:
        html_content = response.content
    
  • 降低请求频率:频繁请求会触发反爬机制,在请求之间添加延迟:

    import time
    
    # 请求前或请求后添加延迟
    time.sleep(2)
    html_content = fetch_content(url)
    

内容的提问来源于stack exchange,提问作者clint

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 08:25:03