You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup多URL爬取异常:单URL正常多URL静默失败

解决多URL批量爬取时的静默失败问题

看起来你遇到的问题是单URL爬取正常,但多URL批量处理时无输出(静默失败),结合你的代码,我帮你分析下可能的原因和解决方案:

可能的问题根源

  1. 部分URL的页面无目标数据:如果某个赛事链接的页面里不存在main col-md-4 eventData类的div,g_data会是空列表,循环不执行就不会有输出,看起来像“失败”。
  2. 反爬机制拦截:连续请求多个URL时,网站可能会限制你的请求,导致页面返回空或错误,但你的代码没有处理请求异常,所以没有报错提示。
  3. 请求未成功但无反馈:makesoup函数里没有检查请求状态码,即使请求失败(比如403、500),也会继续解析空文本,导致后续无数据输出。

修复后的完整代码

我在你的代码基础上添加了异常处理、状态检查、延迟机制和调试输出,确保每个步骤都有反馈:

import requests
from bs4 import BeautifulSoup
import time

def makesoup(url):
    cookies = {'mycountries' : '101,28,3,102,42,10,18,4,2'}
    try:
        # 添加超时,避免脚本卡住
        r = requests.post(url, cookies=cookies, timeout=10)
        # 检查请求是否成功,失败则抛出异常
        r.raise_for_status()
        return BeautifulSoup(r.text,"lxml")
    except requests.exceptions.RequestException as e:
        print(f"❌ 请求 {url} 失败: {str(e)}")
        return None

def linkscrape(links):
    baseurl = "https://sport-tv-guide.live"
    urllist = []
    
    # 先收集所有链接并打印,确认是否正确抓取
    for link in links:
        finalurl = baseurl + link['href']
        urllist.append(finalurl)
        print(f"✅ 已收集赛事链接: {finalurl}")
    
    print(f"\n开始批量爬取,共 {len(urllist)} 个链接\n")
    
    for idx, singleurl in enumerate(urllist, 1):
        print(f"正在处理第 {idx} 个链接: {singleurl}")
        # 添加2秒延迟,避免触发网站反爬机制
        time.sleep(2)
        
        soup2 = makesoup(url=singleurl)
        if not soup2:
            print(f"⚠️ 跳过无效页面: {singleurl}\n")
            continue
        
        g_data = soup2.find_all('div', {'class': 'main col-md-4 eventData'})
        if not g_data:
            print(f"ℹ️ {singleurl} 未找到赛事数据\n")
            continue
        
        # 解析并打印赛事信息
        for match in g_data:
            hometeam = match.find('div', class_='cell40 text-center teamName1').text.strip()
            awayteam = match.find('div', class_='cell40 text-center teamName2').text.strip()
            dateandtime = match.find('div', class_='timeInfo').text.strip()
            
            print(f"🏆 Match: {hometeam} vs {awayteam}")
            print(f"⏰ Date and Time: {dateandtime}\n")

def matches():
    print("🔍 正在获取网球赛事主页面...")
    soup = makesoup(url = "https://sport-tv-guide.live/live/tennis")
    if not soup:
        print("❌ 获取主页面失败,请检查网络或Cookie设置")
        return
    
    links = soup.find_all('a', {'class': 'article flag', 'href' : True})
    print(f"✅ 主页面解析完成,找到 {len(links)} 个赛事链接\n")
    
    if not links:
        print("ℹ️ 未找到任何赛事链接")
        return
    
    linkscrape(links=links)

# 执行主函数
if __name__ == "__main__":
    matches()

关键改进点

  • 异常处理:捕获请求过程中的所有异常(超时、连接失败、状态码错误),并打印明确的错误信息,避免静默失败。
  • 调试输出:每个步骤都有对应的提示(收集链接、处理进度、数据状态),让你清楚知道脚本在做什么。
  • 反爬规避:添加time.sleep(2)延迟,降低请求频率,减少被网站拦截的概率。
  • 空值检查:在解析页面前后,判断soup和g_data是否为空,避免因无数据导致的无输出。

现在运行这个修改后的脚本,你就能看到每个链接的处理状态,即使某个链接没有数据或者请求失败,也会有明确的提示,不会再出现“静默失败”的情况。

内容的提问来源于stack exchange,提问作者Brendan Rodgers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 15:47:55