为什么Selenium循环爬取数据时仅重复返回第一个元素的内容?
问题根源
- 循环内所有元素查找都调用了全局
driver对象的查找方法,默认匹配整个页面的第一个符合条件的元素,所以每次循环拿到的都是第一首歌的数据 - XPath路径以
//开头代表全局搜索,哪怕是基于单个元素查找也会忽略当前上下文,要查找当前元素的子元素需要把XPath开头改为.,代表相对当前节点搜索 - 播放量的XPath硬编码了
li[1],强制固定取列表第一个歌曲的播放量数据,完全没有适配循环逻辑 - 循环变量
song在循环内被字典赋值覆盖,属于不规范写法,容易引发后续逻辑异常
修正后代码
import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd ser= Service(r"C:\Program Files (x86)\chromedriver.exe") options = webdriver.ChromeOptions() options.add_experimental_option('excludeSwitches', ['enable-logging']) driver = webdriver.Chrome(options=options,service=ser) driver.get('https://soundcloud.com/jujubucks') print(driver.title) # 显式等待歌曲列表加载完成 WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CLASS_NAME, 'soundList__item')) ) # 可选:滚动页面加载更多懒加载内容,可根据需要调整滚动次数 for _ in range(3): driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) song_contents = driver.find_elements(By.CLASS_NAME, 'soundList__item') song_list = [] for song in song_contents: # 全部改为从当前song元素下查找子元素,XPath加.改为相对路径 search = song.find_element(By.CLASS_NAME, 'soundTitle__usernameText').text search_song = song.find_element(By.XPATH, './/span[@class=""]').text search_date = song.find_element(By.CLASS_NAME, 'sc-visuallyhidden').text # 播放量改为相对路径查找,用类名匹配更稳定,去掉硬编码的li[1] search_plays = song.find_element(By.XPATH, './/li[contains(@class,"sc-ministats-plays")]/span/span[2]').text song_item ={ 'Artist': search, 'Song_title': search_song, 'Date': search_date, 'Streams': search_plays } song_list.append(song_item) df = pd.DataFrame(song_list) print(df) driver.quit()
额外说明
- Windows路径前加了
r避免转义字符识别错误 - 补充了显式等待和懒加载滚动逻辑,解决SoundCloud动态加载内容拿不全的问题
- 播放量的XPath改为用类名匹配,不会因为页面结构微调就失效,稳定性更强
内容的提问来源于stack exchange,提问作者Houston Khanyile
相关产品推荐
相关产品推荐

