You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Discord机器人绕过SteamDB浏览器检查报错解决

问题根源

你遇到的IndexError本质是请求被SteamDB的Cloudflare反爬机制拦截了,仅配置老旧的Chrome 58版本User-Agent根本无法通过校验,requests拿到的返回内容是拦截提示页,不存在#table-sortable表格元素,选择器匹配不到内容自然会抛出索引越界错误。
另外你的现有代码还有两个逻辑bug:

  • 请求应用详情页后没有重新解析新响应的HTML,直接复用搜索页的soup对象提取详情字段,就算绕过反爬也无法正常拿到数据
  • 月份替换逻辑中str.replace()不会修改原字符串,没有接收返回值会导致替换不生效
  • 请求完成后加的10秒延迟完全无意义,反爬校验发生在请求阶段,拿到响应后再等待不会改变返回内容
可行实现方案
  • 轻量方案:使用专门适配Cloudflare反爬的请求库替换原生requests,自动模拟真实浏览器的TLS指纹、完成JS挑战,不需要启动完整浏览器,资源占用极低。
  • 高通过率方案:使用无头浏览器工具(Playwright/Selenium)启动真实的Chromium内核,隐藏自动化控制特征,等待页面JS渲染完成后再提取内容,适配绝大多数反爬场景,缺点是资源占用比纯请求库高。
  • 最优方案:直接调用Steam官方公开的Web API查询应用信息,没有反爬限制,稳定性远高于爬取第三方页面,不需要处理任何校验逻辑。
代码修复示例

先安装依赖:
pip install cloudscraper beautifulsoup4
核心修复后的代码:

import asyncio
import cloudscraper
from bs4 import BeautifulSoup

async def steam(ctx, *, search):
    await ctx.send('**Шукаю...**')
    # 初始化过反爬的请求客户端,自动模拟最新版Chrome指纹
    scraper = cloudscraper.create_scraper(
        browser={
            'browser': 'chrome',
            'platform': 'windows',
            'desktop': True
        }
    )
    search_keyword = search.replace(" ", "+")
    # 请求搜索页
    search_res = scraper.get(
        f'https://steamdb.info/search/?a=app&q={search_keyword}&type=1&category=0',
        timeout=15
    )
    search_soup = BeautifulSoup(search_res.text, 'html.parser')
    # 先校验元素是否存在,避免拦截时抛错
    first_result = search_soup.select('#table-sortable tr:nth-child(1) td:nth-child(1) a')
    if not first_result:
        await ctx.send('Помилка доступу або результатів не знайдено')
        return
    print(first_result[0].getText().strip())

    app_path = ''
    for url in search_soup.find_all('a', href=True):
        if "/app/" in url['href']:
            app_path = url['href']
            break
    if not app_path:
        await ctx.send('Не вдалося знайти сторінку додатку')
        return
    
    app_url = 'https://steamdb.info' + app_path
    # 请求详情页后重新解析HTML
    app_res = scraper.get(app_url, timeout=15)
    app_soup = BeautifulSoup(app_res.text, 'html.parser')
    print(app_url)

    name = app_soup.find(itemprop='name').getText()
    desc = app_soup.find(itemprop='description').getText()
    reldate_elem = app_soup.select('table.table.table-bordered tr:nth-child(8) td:nth-child(2)')
    if not reldate_elem:
        reldate = 'Невідома дата релізу'
    else:
        reldate = reldate_elem[0].getText().strip()
        reldate_mo = ['January', 'February', 'March', 'April', 'May', 'June', 'July', 'August', 'September', 'October', 'November', 'December']
        reldate_mo_t = ['січ', 'лют', 'бер', 'кві', 'тра', 'чер', 'лип', 'сер', 'вер', 'жов', 'лист', 'гру']
        reldate = reldate.split(' –')[0]
        for i in range(len(reldate_mo)):
            if reldate_mo[i] in reldate:
                # 修复字符串替换不生效的bug
                reldate = reldate.replace(reldate_mo[i], reldate_mo_t[i])
    # 后续拼接消息发送的逻辑自行补充即可
注意事项
  • 控制请求频率,两次请求之间加2秒以上的随机延迟,不要短时间发起大量请求,否则依然会被IP封禁
  • 如果轻量请求库触发高级校验,切换为无头浏览器方案,启动时添加隐藏自动化特征的参数即可正常访问
  • 个人非商用使用请遵守站点规则,不要对站点服务造成压力

内容的提问来源于stack exchange,提问作者enderport

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 17:51:22