You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取分页网站时持续出现500 Internal Server Error求助

爬虫500 Internal Server Error 解决思路

针对你爬取该分页网站时遇到的500错误,给出以下具体排查和修复步骤:

  • 添加请求状态码校验,精准定位问题
    你的代码直接调用requests.get().content,无法直观确认请求状态。先捕获响应对象,判断状态码,明确是特定页面还是所有页面都返回错误:

    import requests
    import pandas as pd
    from bs4 import BeautifulSoup
    from random import randint
    import time
    
    url = "https://www.linguasport.com/futbol/nacional/liga/seekff_esp.asp?pn={}"
    headers = {'User-Agent': "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
    
    dfs = []
    for page in range(1, 20):
        time.sleep(randint(1,5))
        response = requests.get(url.format(page), headers=headers)
        # 新增状态码校验逻辑
        if response.status_code != 200:
            print(f"第{page}页请求失败,状态码:{response.status_code}")
            continue
        soup = BeautifulSoup(response.content, "html.parser")
        tds = soup.find_all("td")
        # 后续处理td数据...
    
  • 补充完整请求头,模拟真实浏览器行为
    仅携带User-Agent可能被服务器识别为爬虫,补充浏览器请求中的Referer、Accept-Language等字段(可从浏览器开发者工具中复制完整请求头):

    headers = {
        'User-Agent': "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
        'Referer': "https://www.linguasport.com/futbol/nacional/liga/seekff_esp.asp?pn=1",
        'Accept-Language': "zh-CN,zh;q=0.9",
        'Accept': "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"
    }
    
  • 使用会话维持Cookie,保证请求合法性
    部分网站会通过Cookie验证会话有效性,首次访问首页生成的Cookie需要在后续分页请求中携带。用requests.Session()自动维护会话:

    # 初始化会话对象
    session = requests.Session()
    session.headers.update(headers)
    # 先访问首页获取有效Cookie
    session.get("https://www.linguasport.com/futbol/nacional/liga/seekff_esp.asp?pn=1")
    
    dfs = []
    for page in range(1, 20):
        time.sleep(randint(1,5))
        response = session.get(url.format(page))
        if response.status_code != 200:
            print(f"第{page}页请求失败,状态码:{response.status_code}")
            continue
        soup = BeautifulSoup(response.content, "html.parser")
        tds = soup.find_all("td")
        # 后续处理td数据...
    
  • 验证分页范围的有效性
    手动访问网站确认实际最大页码,若你设置的range(1,20)超过网站实际分页数量,请求不存在的页面会返回500错误。先缩小范围测试,比如先爬1-5页,确认正常后再扩展。

  • 修复代码遗漏的导入
    你的代码中使用了randint但未导入random模块,且sleep需调用time.sleep,需补充:

    from random import randint
    import time
    

内容的提问来源于stack exchange,提问作者console.log

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 09:05:44