You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取图片时IndexError异常捕获失效问题

问题分析与解决建议

我看了你的代码和报错信息,核心问题出在你把触发IndexError的代码放在了try-except块外面,导致异常根本没被捕获到!

具体来说,这行代码:

img_src = img_page[0].get('src', '')

在进入else分支的try块之前就执行了。当soup.find('div', {'itemprop' : 'blogPost'}).find_all('img')返回空列表时,访问索引0直接抛出IndexError,而你的try-except只包裹了后面的判断逻辑,完全没机会处理这个错误。

修正方案:重构逻辑,优先判断而非依赖异常捕获

我推荐用更清晰的逻辑来处理“找图”流程,避免嵌套过深,同时提前判断列表是否为空,而不是靠try-except兜底(当然用try-except也可以,但判断空列表更直观):

site_link = []
site_img = []
for i in site_links:
    try:
        # 先处理请求,捕获可能的网络异常
        r = requests.get(i, timeout=10).text
        soup = bs4.BeautifulSoup(r, 'html5lib')
        image_found = False
        
        # 1. 优先检查画廊图片
        img_gallery = soup.find('a', {'class':'sigProLink fancybox-gallery', 'href':True})
        if img_gallery:
            href = img_gallery.get('href', '')
            if '.jpg' in href:
                img_link = '***GALLERY*** ' + href
                site_img.append(img_link)
                print(img_link)
                image_found = True
        
        # 2. 画廊没找到,检查页面内图片
        if not image_found:
            img_page = soup.find('div', {'itemprop' : 'blogPost'}).find_all('img')
            # 先判断列表是否非空,再访问索引0
            if img_page:
                img_src = img_page[0].get('src', '')
                if '.jpg' in img_src:
                    img_link = '**PAGE*** ' + img_src
                    site_img.append(img_link)
                    print(img_link)
                    image_found = True
        
        # 3. 两种位置都没找到图片
        if not image_found:
            site_img.append('No Images Found')
            print('No Images Found')
    
    # 捕获所有可能的异常(网络错误、解析错误等),确保脚本不中断
    except Exception as e:
        print(f"处理URL {i}时出错: {str(e)}")
        site_img.append('Error Processing Page')

关键改进点:

  1. 把请求和解析逻辑整体包裹在try-except里:不仅能处理IndexError,还能捕获网络连接超时、页面解析失败等其他异常,保证82个URL的遍历不会中途中断。
  2. 用image_found标志位跟踪状态:避免多层嵌套的if-else,逻辑更清晰。
  3. 提前判断列表是否为空:比如if img_page:代替直接访问img_page[0],从根源避免IndexError。
  4. 添加请求超时:requests.get(i, timeout=10)防止某个慢响应的URL卡住整个脚本。

可选优化:

如果想保留try-except的写法,只需要把可能触发IndexError的代码移到try块内即可:

# 替换原来的else分支逻辑
else:
    try:
        img_page = soup.find('div', {'itemprop' : 'blogPost'}).find_all('img')
        img_src = img_page[0].get('src', '')
        if '.jpg' in img_src:
            img_link = '**PAGE*** ' + img_src
            site_img.append(img_link)
            print(img_link)
    except IndexError:
        site_img.append('No Images Found')
        print('No Images Found')

不过这种写法不如前面的标志位方案直观,而且没法处理其他潜在异常。

内容的提问来源于stack exchange,提问作者prime90

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:46:43