You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何爬取Spotify播放列表数据?代码返回空列表求解答

解决Spotify播放列表爬取空列表问题

问题原因

Spotify网页采用动态渲染机制,直接用requests获取的静态HTML里没有实际的歌曲数据——所有曲目信息都是通过JavaScript在浏览器加载完成后才生成的,所以BeautifulSoup根本找不到目标元素。

解决方案一:使用Spotify官方API(推荐)

Spotify提供了官方API来获取播放列表数据,这是最稳定、合规的方式。

实现步骤:

  1. 前往Spotify开发者平台创建应用,获取专属的client_id和client_secret。
  2. 通过客户端凭证流获取访问令牌。
  3. 调用API接口获取播放列表曲目详情。

示例代码:

import requests

# 替换为你的client_id和client_secret
CLIENT_ID = "your_client_id"
CLIENT_SECRET = "your_client_secret"
PLAYLIST_ID = "37i9dQZEVXbNG2KDcFcKOF"

# 获取访问令牌
auth_url = "https://accounts.spotify.com/api/token"
auth_response = requests.post(auth_url, {
    'grant_type': 'client_credentials',
    'client_id': CLIENT_ID,
    'client_secret': CLIENT_SECRET,
})
auth_data = auth_response.json()
access_token = auth_data['access_token']

# 获取播放列表曲目
headers = {
    'Authorization': f'Bearer {access_token}'
}
playlist_url = f"https://api.spotify.com/v1/playlists/{PLAYLIST_ID}/tracks"
response = requests.get(playlist_url, headers=headers)
tracks = response.json()['items']

# 提取歌曲名与艺术家信息
data = []
for track in tracks:
    song_name = track['track']['name']
    artist_name = track['track']['artists'][0]['name']
    data.append([song_name, artist_name])
    print([song_name, artist_name])

解决方案二:使用Selenium模拟浏览器加载动态内容

如果不想依赖API,可以用Selenium模拟浏览器打开页面,等待JavaScript加载完成后再解析HTML。

示例代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

URL = "https://open.spotify.com/playlist/37i9dQZEVXbNG2KDcFcKOF"

# 初始化浏览器(需提前下载对应浏览器的驱动,比如ChromeDriver)
driver = webdriver.Chrome()
driver.get(URL)

# 等待曲目列表加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "div[data-testid='tracklist-row']"))
)

# 获取渲染后的页面源码并关闭浏览器
page_source = driver.page_source
driver.quit()

# 解析页面提取数据
soup = BeautifulSoup(page_source, "lxml")
song_rows = soup.find_all("div", {"data-testid": "tracklist-row"})

data = []
for row in song_rows:
    song_name = row.find("span", attrs={"data-encore-id": "type"}).text
    artist_name = row.find_all("span", attrs={"data-encore-id": "type"})[1].text
    data.append([song_name, artist_name])
    print([song_name, artist_name])

注意事项:

  • 使用Selenium需确保浏览器驱动版本与浏览器版本匹配。
  • 频繁爬取可能触发Spotify反爬机制,注意控制请求频率。

内容的提问来源于stack exchange,提问作者Santhiya s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 15:07:15