You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为什么Selenium仅迭代5次就停止,无法完成完整数据集的爬取?

问题根因

  • 你爬取的SoundCloud页面采用懒加载机制:首次打开页面时仅渲染前5条作品数据,剩余内容需要向下滚动页面触发加载请求后才会插入到DOM结构中。你的代码在点击cookie同意按钮后立刻调用find_elements获取元素,此时拿到的song_contents列表本身就只有5个节点,所以循环仅执行5次就结束了。
  • 额外潜在问题:如果滚动过程中DOM发生刷新,提前获取的元素列表会触发StaleElementReferenceException异常,直接终止遍历逻辑。

解决方法

你需要先模拟滚动操作加载所有作品数据,再统一提取内容,步骤如下:

  1. 处理完cookie同意弹窗后,循环执行滚动到页面底部→等待新内容加载的操作,直到页面高度不再变化,确认所有内容加载完成
  2. 加载完成后再一次性获取所有作品元素,遍历提取数据

修改后可运行代码

import time
import random
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

ser= Service("C:\Program Files (x86)\chromedriver.exe")
options = webdriver.ChromeOptions()
options.add_experimental_option('excludeSwitches', ['enable-logging'])
driver = webdriver.Chrome(options=options,service=ser)
driver.get('https://soundcloud.com/jujubucks')
print(driver.title)

wait = WebDriverWait(driver,30)
# 处理cookie同意
wait.until(EC.element_to_be_clickable((By.ID,"onetrust-accept-btn-handler"))).click()

# 新增滚动加载全量内容逻辑
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 随机等待2-4秒,避免加载不完整同时降低反爬检测风险
    time.sleep(random.uniform(2,4))
    new_height = driver.execute_script("return document.body.scrollHeight")
    if new_height == last_height:
        break
    last_height = new_height

# 全量加载完成后再获取元素
song_contents = driver.find_elements(By.CLASS_NAME, 'soundList__item')
song_list = []

for option in song_contents:
    search = option.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__username')]/span").text
    search_song = option.find_element(By.XPATH, ".//a[contains(@class,'soundTitle__title')]/span").text
    search_date = option.find_element(By.XPATH, ".//time[contains(@class,'relativeTime')]/span").text
    search_plays = option.find_element(By.XPATH, ".//span[contains(@class,'sc-ministats-small')]/span").text
    
    item ={
        'Artist': search, 
        'Song_title': search_song, 
        'Date': search_date,
        'Streams': search_plays
    }
    song_list.append(item)

df = pd.DataFrame(song_list)
print(df)
print(f"共获取到{len(song_list)}条数据")

driver.quit()

注意事项

  • 如果页面作品数量过多,可以适当增加滚动后的等待时长,避免网络延迟导致新内容还未渲染就判断停止加载
  • 可以根据需求调整滚动逻辑,比如设置最大滚动次数,避免页面无限加载导致程序卡死

内容的提问来源于stack exchange,提问作者Houston Khanyile

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 01:27:01