You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask+BeautifulSoup构建API本地正常,Heroku部署后网页爬取失效问题求助

解决Heroku上Flask爬虫失效的问题

嘿,你猜的没错——Heroku上爬虫失效大概率和请求头(尤其是User-Agent)被目标网站拦截有关,另外还有几个Heroku特有的坑也可能导致这个问题,我给你整理几个实用的解决方案:

1. 完善请求头,模拟真实浏览器

很多网站会直接拦截requests库默认的User-Agent(比如python-requests/2.28.1),尤其是Heroku这类云平台的IP更容易被标记成机器人。你需要把请求头改成和真实浏览器一模一样的:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://目标网站的主页地址/',  # 比如你爬的是体育比分站,就填它的主页
    'DNT': '1',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

把这个替换你原来的headers,尽量覆盖浏览器发送的核心字段,降低被识别的概率。

2. 细化错误处理,排查具体问题

你现在的except太宽泛了,根本不知道是请求被拦截了、页面结构变了还是网络超时。修改代码,捕获具体异常并打印日志,方便在Heroku上排查:

try:
    # 添加上超时设置,避免Heroku因请求超时终止进程
    page = requests.get(teem, headers=headers, timeout=10)
    # 主动抛出HTTP错误(比如403、500),方便定位
    page.raise_for_status()
    soup = BeautifulSoup(page.content, 'html.parser')
    score_element = soup.find(class_="imso_mh__wl imso-ani imso_mh__tas")
    # 先判断元素是否存在,避免空指针错误
    if not score_element:
        return 'ERROR: 未找到比分元素'
    score = score_element.get_text()
    return translator.translate(score.strip(), dest="en").text
except requests.exceptions.RequestException as e:
    # 把错误信息打印到Heroku日志
    print(f"请求出错: {str(e)}")
    return f'ERROR: 请求失败 - {str(e)}'
except Exception as e:
    print(f"意外错误: {str(e)}")
    return 'ERROR: 未知问题'

然后用Heroku命令查看实时日志:

heroku logs --tail

如果日志里出现403 Forbidden,那就是反爬拦截;如果是元素找不到,可能目标网站在Heroku环境下返回的是移动端页面,结构和本地不一样。

3. 处理Heroku IP被拉黑的情况

Heroku的共享IP池可能已经被目标网站加入黑名单了,这时候光改User-Agent没用。可以试试:

  • 用付费代理IP(比如BrightData、Oxylabs),把代理配置到requests.get里:
    proxies = {
        'http': 'http://你的代理地址',
        'https': 'https://你的代理地址'
    }
    page = requests.get(teem, headers=headers, proxies=proxies, timeout=10)
    
  • 如果目标网站有官方API,优先用API替代爬取,既稳定又合规,完全避免反爬问题。

4. 应对Cloudflare等反爬系统

如果目标网站用了Cloudflare,普通的requests根本绕不过去。可以用cloudscraper库代替requests,它能自动处理Cloudflare的验证:

  1. 在requirements.txt里添加cloudscraper
  2. 修改代码:
    import cloudscraper
    scraper = cloudscraper.create_scraper()
    page = scraper.get(teem, headers=headers, timeout=10)
    

按上面的步骤排查,应该能解决Heroku上的爬虫问题。

内容的提问来源于stack exchange,提问作者Sahaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:18:12