为什么Selenium仅迭代5次就停止,无法完成完整数据集的爬取?
问题根因
- 你爬取的SoundCloud页面采用懒加载机制:首次打开页面时仅渲染前5条作品数据,剩余内容需要向下滚动页面触发加载请求后才会插入到DOM结构中。你的代码在点击cookie同意按钮后立刻调用
find_elements获取元素,此时拿到的song_contents列表本身就只有5个节点,所以循环仅执行5次就结束了。 - 额外潜在问题:如果滚动过程中DOM发生刷新,提前获取的元素列表会触发
StaleElementReferenceException异常,直接终止遍历逻辑。
解决方法
你需要先模拟滚动操作加载所有作品数据,再统一提取内容,步骤如下:
- 处理完cookie同意弹窗后,循环执行滚动到页面底部→等待新内容加载的操作,直到页面高度不再变化,确认所有内容加载完成
- 加载完成后再一次性获取所有作品元素,遍历提取数据
修改后可运行代码
import time import random from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd ser= Service("C:\Program Files (x86)\chromedriver.exe") options = webdriver.ChromeOptions() options.add_experimental_option('excludeSwitches', ['enable-logging']) driver = webdriver.Chrome(options=options,service=ser) driver.get('https://soundcloud.com/jujubucks') print(driver.title) wait = WebDriverWait(driver,30) # 处理cookie同意 wait.until(EC.element_to_be_clickable((By.ID,"onetrust-accept-btn-handler"))).click() # 新增滚动加载全量内容逻辑 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 随机等待2-4秒,避免加载不完整同时降低反爬检测风险 time.sleep(random.uniform(2,4)) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 全量加载完成后再获取元素 song_contents = driver.find_elements(By.CLASS_NAME, 'soundList__item') song_list = [] for option in song_contents: search = option.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__username')]/span").text search_song = option.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__title')]/span").text search_date = option.find_element(By.XPATH, ".//time[contains(@class,'relativeTime')]/span").text search_plays = option.find_element(By.XPATH, ".//span[contains(@class,'sc-ministats-small')]/span").text item ={ 'Artist': search, 'Song_title': search_song, 'Date': search_date, 'Streams': search_plays } song_list.append(item) df = pd.DataFrame(song_list) print(df) print(f"共获取到{len(song_list)}条数据") driver.quit()
注意事项
- 如果页面作品数量过多,可以适当增加滚动后的等待时长,避免网络延迟导致新内容还未渲染就判断停止加载
- 可以根据需求调整滚动逻辑,比如设置最大滚动次数,避免页面无限加载导致程序卡死
内容的提问来源于stack exchange,提问作者Houston Khanyile
相关产品推荐
相关产品推荐

