You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium获取Spotify艺人href返回异常值问题求助

问题分析与解决

核心错误原因

你代码里的关键错误是循环中误将links列表本身追加到了links列表中,而非获取到的link变量:

links.append(links)  # 错误写法

这会导致links列表不断嵌套自身,最终生成[[...], [...], ...]这种递归结构,写入CSV后就呈现出你看到的异常内容。

修复后的代码

基础修复(修正追加逻辑)

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
import pandas as pd

website= "https://open.spotify.com/playlist/6iwz7yurUKaILuykiyeztu"
path= "C:/Users/ashut/Downloads/Misc Docs/chromedriver_win32/chromedriver.exe"

service=Service(executable_path=path)
driver=webdriver.Chrome(service=service)

driver.get(website) 

containers=driver.find_elements(by="xpath",value='//div[@data-testid="tracklist-row"]/div[@aria-colindex="2"]/div')

titles = []
artists = []
links = []

for container in containers:
    title=container.find_element(by="xpath", value='./a/div').text
    artist=container.find_element(by="xpath", value='./span/a').text
    link=container.find_element(by="xpath", value='./span/a').get_attribute("href")
    titles.append(title)
    artists.append(artist)
    links.append(link)  # 修正为追加获取到的link变量
    
mydict={'titles':titles,'artists':artists,'links':links}
artistslist= pd.DataFrame(mydict)
artistslist.to_csv('list_of_artist.csv', index=False)  # 可选:移除CSV默认生成的索引列

优化建议(处理动态加载)

Spotify是动态渲染页面,可能存在元素未加载完成就执行查找的情况,添加显式等待可提升爬取稳定性:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import pandas as pd

website= "https://open.spotify.com/playlist/6iwz7yurUKaILuykiyeztu"
path= "C:/Users/ashut/Downloads/Misc Docs/chromedriver_win32/chromedriver.exe"

service=Service(executable_path=path)
driver=webdriver.Chrome(service=service)
wait = WebDriverWait(driver, 10)  # 最多等待10秒直至元素加载

driver.get(website) 

# 等待曲目列表加载完成后再执行后续操作
wait.until(EC.presence_of_all_elements_located((By.XPATH, '//div[@data-testid="tracklist-row"]')))

containers=driver.find_elements(by="xpath",value='//div[@data-testid="tracklist-row"]/div[@aria-colindex="2"]/div')

titles = []
artists = []
links = []

for container in containers:
    title=container.find_element(by="xpath", value='./a/div').text
    artist=container.find_element(by="xpath", value='./span/a').text
    link=container.find_element(by="xpath", value='./span/a').get_attribute("href")
    titles.append(title)
    artists.append(artist)
    links.append(link)
    
mydict={'titles':titles,'artists':artists,'links':links}
artistslist= pd.DataFrame(mydict)
artistslist.to_csv('list_of_artist.csv', index=False)

修复后效果

修正后,CSV的links列会正确显示艺人页面的完整URL,不再出现嵌套列表的异常内容。


内容的提问来源于stack exchange,提问作者Ashuwathama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 21:57:35