You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取Spotify播放列表仅获前20条,求解决方案

Spotify播放列表爬取:无法获取完整歌曲列表

问题现象

  • 仅能获取页面初始加载的歌曲,滚动、等待操作无法加载全部内容
  • 缩小浏览器窗口仅能多获取20-30条结果
  • 手动滚动后爬取会跳过前几首歌曲,直接从当前已加载部分开始

原代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
import pandas as pd
import time
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

website= "https://open.spotify.com/playlist/6iwz7yurUKaILuykiyeztu"
path= "C:/Users/ashut/Downloads/Misc Docs/chromedriver_win32/chromedriver.exe"

service=Service(executable_path=path)
driver=webdriver.Chrome(service=service)

driver.get(website) 
containers=driver.find_elements(by="xpath",value='//div[@data-testid="tracklist-row"]/div[@aria-colindex="2"]/div')

titles = []
artists = []
links = []

for container in containers:
    title=container.find_element(by="xpath", value='./a/div').text
    artist=container.find_element(by="xpath", value='./span/a').text
    link=container.find_element(by="xpath", value='./span/a').get_attribute("href")
    titles.append(title)
    artists.append(artist)
    links.append(link)
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)
    
mydict={'titles':titles,'artists':artists,'links':links}
artistslist= pd.DataFrame(mydict)
artistslist.to_csv('list_of_artist.csv')

问题分析

  1. 初始容器抓取过早:代码一开始就获取所有containers,但此时页面仅加载了部分歌曲,后续滚动加载的新内容不会被加入这个静态列表
  2. 滚动逻辑不匹配:滚动到页面底部的方式无法触发Spotify的懒加载机制,且滚动时机错误,无法持续加载新歌曲
  3. 无去重处理:重新抓取容器时会重复获取已爬取的内容,导致数据冗余

修正后的代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
import pandas as pd
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import time

website = "https://open.spotify.com/playlist/6iwz7yurUKaILuykiyeztu"
path = "C:/Users/ashut/Downloads/Misc Docs/chromedriver_win32/chromedriver.exe"

service = Service(executable_path=path)
driver = webdriver.Chrome(service=service)
driver.get(website)
driver.maximize_window()

# 等待页面初始加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, '//div[@data-testid="tracklist-row"]'))
)

track_data = {}
last_count = 0

while True:
    # 获取当前所有歌曲容器
    containers = driver.find_elements(by="xpath", value='//div[@data-testid="tracklist-row"]/div[@aria-colindex="2"]/div')
    
    # 遍历容器,跳过已爬取内容
    for container in containers:
        try:
            link = container.find_element(by="xpath", value='./span/a').get_attribute("href")
            if link not in track_data:
                title = container.find_element(by="xpath", value='./a/div').text
                artist = container.find_element(by="xpath", value='./span/a').text
                track_data[link] = {"title": title, "artist": artist}
        except Exception:
            # 跳过加载异常的条目
            continue
    
    # 判断是否还有新内容加载
    current_count = len(track_data)
    if current_count == last_count:
        break
    last_count = current_count
    
    # 滚动到最后一条歌曲,触发懒加载
    last_track = containers[-1]
    driver.execute_script("arguments[0].scrollIntoView({behavior: 'smooth', block: 'center'});", last_track)
    time.sleep(2)  # 可根据网络情况调整等待时间

# 转换为DataFrame并保存
artistslist = pd.DataFrame.from_dict(track_data, orient='index').reset_index(drop=True)
artistslist = artistslist[["title", "artist"]]
artistslist.to_csv('list_of_artist.csv', index=False)

driver.quit()

核心优化说明

  • 循环加载直到无新内容:通过对比每次爬取的歌曲数量,判断是否加载完成
  • 精准滚动触发懒加载:滚动到最后一条已加载歌曲的位置,适配Spotify的加载逻辑,比滚动到页面底部更可靠
  • 去重机制:用歌曲链接作为唯一标识存储数据,避免重复抓取
  • 等待逻辑优化:先显式等待页面初始加载,滚动后用固定等待让新内容加载完成,平衡效率与稳定性

内容的提问来源于stack exchange,提问作者Ashuwathama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 00:54:33