You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用for循环遍历列表时如何跳过触发异常的项?Selenium爬虫场景

问题原因分析
  • 异常捕获范围错误:你把整个for循环包裹在全局try块中,只要任意一首歌曲的元素查找触发NoSuchElementException,就会直接跳出循环执行全局except的pass逻辑,后续所有列表项都不会被处理。
  • 判断逻辑无效:if search_plays == False 完全不会生效,一是因为找不到播放量元素时会直接抛异常,代码根本执行不到这一行;二是就算成功拿到search_plays的值,它是字符串类型,和布尔值False做恒等判断永远返回False,等于没写这个判断。
  • continue使用限制的问题:你之前把try放在循环外面,当然不能在except里用continue,只要把try-except放到循环体内部,捕获当前项的异常后直接用continue就完全合法。
修复后的代码
import time
from selenium import webdriver
import selenium
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions import EC
import pandas as pd

ser= Service("C:\Program Files (x86)\chromedriver.exe")
options = webdriver.ChromeOptions()
options.add_experimental_option('excludeSwitches', ['enable-logging'])
driver = webdriver.Chrome(options=options,service=ser)
driver.get('https://soundcloud.com/jujubucks')
print(driver.title)

wait = WebDriverWait(driver,30)
wait.until(EC.element_to_be_clickable((By.ID,"onetrust-accept-btn-handler"))).click()

# 歌曲列表移到循环外初始化,避免每次覆盖之前的结果
song_list = []
i = 1
for _ in range(32):
    # try-except放到循环内部,仅捕获当前单项的处理异常
    try:
        song_contents = driver.find_element(By.XPATH, "//li[@class='soundList__item'][{}]".format(i))
        driver.execute_script("arguments[0].scrollIntoView(true);",song_contents)
        search = song_contents.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__username')]/span").text
        search_song = song_contents.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__title')]/span").text
        search_date = song_contents.find_element(By.XPATH, ".//time[contains(@class,'relativeTime')]/span").text
        search_plays = song_contents.find_element(By.XPATH, ".//span[contains(@class,'sc-ministats-small')]/span").text
        
        # 修正空值判断逻辑,判断播放量字符串是否为空
        if not search_plays.strip():
            continue
        
        option ={
        'Artist': search, 
        'Song_title': search_song, 
        'Date': search_date,
        'Streams': search_plays
        }
        song_list.append(option)
        print(pd.DataFrame([option]))
    # 精准捕获元素不存在异常,避免吞掉其他错误
    except selenium.common.exceptions.NoSuchElementException:
        # 异常直接跳过当前项,继续下一轮循环
        pass
    finally:
        # 无论当前项处理成功失败,索引都自增,避免卡死在同一个位置
        i +=1

# 循环结束后生成完整结果表
df = pd.DataFrame(song_list)
print("完整爬取结果:")
print(df)

driver.quit()
修复说明
  • 调整了try-except的位置,放到循环体内部,仅捕获单个列表项的处理异常,出错后直接跳过当前项,不会中断整个爬取流程
  • 修正了空播放量的判断逻辑,改为判断字符串是否为空,符合实际业务场景
  • 把song_list初始化移到循环外部,避免每次循环覆盖之前的爬取结果
  • 索引i的增量放到finally块中,不管当前项处理成功失败都会自增,不会卡住重复处理同一个索引的项
  • 缩小了异常捕获范围,仅捕获NoSuchElementException,避免其他未知错误被直接吞掉无法排查

内容的提问来源于stack exchange,提问作者Houston Khanyile

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 05:36:04