Python使用for循环遍历列表时如何跳过触发异常的项?Selenium爬虫场景
问题原因分析
- 异常捕获范围错误:你把整个for循环包裹在全局try块中,只要任意一首歌曲的元素查找触发
NoSuchElementException,就会直接跳出循环执行全局except的pass逻辑,后续所有列表项都不会被处理。 - 判断逻辑无效:
if search_plays == False完全不会生效,一是因为找不到播放量元素时会直接抛异常,代码根本执行不到这一行;二是就算成功拿到search_plays的值,它是字符串类型,和布尔值False做恒等判断永远返回False,等于没写这个判断。 - continue使用限制的问题:你之前把try放在循环外面,当然不能在except里用continue,只要把try-except放到循环体内部,捕获当前项的异常后直接用continue就完全合法。
修复后的代码
import time from selenium import webdriver import selenium from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions import EC import pandas as pd ser= Service("C:\Program Files (x86)\chromedriver.exe") options = webdriver.ChromeOptions() options.add_experimental_option('excludeSwitches', ['enable-logging']) driver = webdriver.Chrome(options=options,service=ser) driver.get('https://soundcloud.com/jujubucks') print(driver.title) wait = WebDriverWait(driver,30) wait.until(EC.element_to_be_clickable((By.ID,"onetrust-accept-btn-handler"))).click() # 歌曲列表移到循环外初始化,避免每次覆盖之前的结果 song_list = [] i = 1 for _ in range(32): # try-except放到循环内部,仅捕获当前单项的处理异常 try: song_contents = driver.find_element(By.XPATH, "//li[@class='soundList__item'][{}]".format(i)) driver.execute_script("arguments[0].scrollIntoView(true);",song_contents) search = song_contents.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__username')]/span").text search_song = song_contents.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__title')]/span").text search_date = song_contents.find_element(By.XPATH, ".//time[contains(@class,'relativeTime')]/span").text search_plays = song_contents.find_element(By.XPATH, ".//span[contains(@class,'sc-ministats-small')]/span").text # 修正空值判断逻辑,判断播放量字符串是否为空 if not search_plays.strip(): continue option ={ 'Artist': search, 'Song_title': search_song, 'Date': search_date, 'Streams': search_plays } song_list.append(option) print(pd.DataFrame([option])) # 精准捕获元素不存在异常,避免吞掉其他错误 except selenium.common.exceptions.NoSuchElementException: # 异常直接跳过当前项,继续下一轮循环 pass finally: # 无论当前项处理成功失败,索引都自增,避免卡死在同一个位置 i +=1 # 循环结束后生成完整结果表 df = pd.DataFrame(song_list) print("完整爬取结果:") print(df) driver.quit()
修复说明
- 调整了try-except的位置,放到循环体内部,仅捕获单个列表项的处理异常,出错后直接跳过当前项,不会中断整个爬取流程
- 修正了空播放量的判断逻辑,改为判断字符串是否为空,符合实际业务场景
- 把song_list初始化移到循环外部,避免每次循环覆盖之前的爬取结果
- 索引i的增量放到finally块中,不管当前项处理成功失败都会自增,不会卡住重复处理同一个索引的项
- 缩小了异常捕获范围,仅捕获
NoSuchElementException,避免其他未知错误被直接吞掉无法排查
内容的提问来源于stack exchange,提问作者Houston Khanyile
相关产品推荐
相关产品推荐

