使用Python、Selenium(WebDriver)提取Spotify播放列表主歌手名并保存至CSV
Got it, let's build out the rest of your script to grab those primary artist names and save them to a CSV. Here's a step-by-step breakdown with full working code:
1. Add Required Imports
First, we'll need extra Selenium tools to handle dynamic page content and Python's built-in csv module for saving the data:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import csv import time
2. Handle Lazy Loading (Critical!)
Spotify loads playlist tracks incrementally as you scroll. We need to keep scrolling to the bottom until all tracks are loaded:
driver = webdriver.Chrome() driver.get('https://open.spotify.com/playlist/4FLeoROn5GT7n2tZq5XB4V?si=5wxAuFwgSPCuyjzUx9DDNw') # Wait for the first track to appear (adjust timeout if your internet is slow) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, 'div[data-testid="tracklist-row"]')) ) # Scroll repeatedly to load all tracks last_scroll_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to the bottom of the page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new tracks to load time.sleep(2) # Check if we've reached the end of the playlist new_scroll_height = driver.execute_script("return document.body.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height
3. Extract Primary Artist Names
Each track's primary artist is the first link in the artist section of the track row. We'll locate all these elements and pull their text:
# Find all primary artist elements (first link in each track's artist section) artist_elements = driver.find_elements(By.CSS_SELECTOR, 'div[data-testid="tracklist-row"] span a:first-of-type') # Extract clean text from each element, skipping empty entries primary_artists = [artist.text.strip() for artist in artist_elements if artist.text.strip()]
4. Save to CSV File
Finally, we'll write the extracted artist names to a CSV file:
# Save the list to a CSV with open('spotify_primary_artists.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) # Write a header row writer.writerow(['Primary Artist']) # Write each artist name as a separate row for artist in primary_artists: writer.writerow([artist])
5. Cleanup
Don't forget to close the browser when the script finishes:
driver.quit()
Quick Tips to Avoid Issues:
- Adjust Timings: If you get errors about missing elements, increase the
WebDriverWaittimeout (e.g., from 10 to 15) or thetime.sleep()duration in the scroll loop. - DOM Changes: Spotify occasionally updates its UI. If the CSS selectors stop working, right-click a track's artist and inspect the element—look for stable
data-testidattributes instead of class names. - Login Walls: Some playlists require a Spotify account. If you hit a login page, add code to automate logging in (use
send_keys()on the email/password fields and click the login button).
内容的提问来源于stack exchange,提问作者Adryro

