使用Selenium爬取Bandcamp发现卡片链接时HTML不更新的问题
问题
我写了一个简单的网页爬虫脚本,想抓取Bandcamp平台发现卡片里的所有超链接,但遇到了重复抓取的问题:脚本能正确获取第1页的8个专辑链接,但第2-4页重复抓取第2页的内容,后续也有类似重复情况。虽然Selenium控制的浏览器页面已经更新(URL也变了),但Selenium无法识别HTML变化,导致重复获取相同内容。我试过延长等待时间、用WebDriverWait等待不同元素加载,都没用。
我的代码如下:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.select import Select browser = webdriver.Chrome() all_links = [] page = 1 url = "https://bandcamp.com/?g=all&s=new&p=0&gn=0&f=digital&w=-1" browser.get(url) while page < 6: page += 1 # wait until discover cards are loaded test = WebDriverWait(browser, 20).until(EC.element_to_be_clickable( (By.XPATH, '//*[@id="discover"]/div[9]/div[2]/div/div[1]/div/table/tbody/tr[1]/td[1]/a/div'))) # scrape hyperlinks for each of the 8 albums shown titles = browser.find_elements(By.CLASS_NAME, "item-title") links = [title.get_attribute('href') for title in titles[-8:]] all_links = all_links + links print(links) # pagination - click through the page buttons as the links are scraped page_nums = browser.find_elements(By.CLASS_NAME, 'item-page') for page_num in page_nums: if page_num.text.isnumeric(): if int(page_num.text) == page: page_num.click() time.sleep(20) # I've tried multiple long wait times as well as WebDriverWaits on different elements to see if the HTML will update, but I haven't seen a positive effect break
输出的重复链接示例:
['https://neubauten.bandcamp.com/album/stimmen-reste-musterhaus-7?from=discover-new', 'https://cirka1.bandcamp.com/album/time?from=discover-new', 'https://futuramusicsound.bandcamp.com/album/yoga-meditation?from=discover-new', 'https://deathsoundbatrecordings.bandcamp.com/album/real-mushrooms-dsbep092?from=discover-new', 'https://riacurley.bandcamp.com/album/take-me-album?from=discover-new', 'https://terracuna.bandcamp.com/album/el-origen-del-viento?from=discover-new', 'https://hyper-music.bandcamp.com/album/hypermusic-vol-4?from=discover-new', 'https://defisis1.bandcamp.com/album/priceless?from=discover-new'] ['https://jarnosalo.bandcamp.com/album/here-lies-ancient-blob?from=discover-new', 'https://andreneitzel.bandcamp.com/album/allegasi-gold-2?from=discover-new', 'https://moonraccoon.bandcamp.com/album/prequels?from=discover-new', 'https://lolivone.bandcamp.com/album/live-at-the-berklee-performance-center?from=discover-new', 'https://nilswrasse.bandcamp.com/album/a-calling-from-the-desert-to-the-sea-original-motion-picture-soundtrack?from=discover-new', 'https://whitereaperaskingride.bandcamp.com/album/asking-for-a-ride?from=discover-new', 'https://collageeffect.bandcamp.com/album/emerald-network?from=discover-new', 'https://foxteethnj.bandcamp.com/album/through-the-blue?from=discover-new'] ['https://jarnosalo.bandcamp.com/album/here-lies-ancient-blob?from=discover-new', 'https://andreneitzel.bandcamp.com/album/allegasi-gold-2?from=discover-new', 'https://moonraccoon.bandcamp.com/album/prequels?from=discover-new', 'https://lolivone.bandcamp.com/album/live-at-the-berklee-performance-center?from=discover-new', 'https://nilswrasse.bandcamp.com/album/a-calling-from-the-desert-to-the-sea-original-motion-picture-soundtrack?from=discover-new', 'https://whitereaperaskingride.bandcamp.com/album/asking-for-a-ride?from=discover-new', 'https://collageeffect.bandcamp.com/album/emerald-network?from=discover-new', 'https://foxteethnj.bandcamp.com/album/through-the-blue?from=discover-new'] ['https://jarnosalo.bandcamp.com/album/here-lies-ancient-blob?from=discover-new', 'https://andreneitzel.bandcamp.com/album/allegasi-gold-2?from=discover-new', 'https://moonraccoon.bandcamp.com/album/prequels?from=discover-new', 'https://lolivone.bandcamp.com/album/live-at-the-berklee-performance-center?from=discover-new', 'https://nilswrasse.bandcamp.com/album/a-calling-from-the-desert-to-the-sea-original-motion-picture-soundtrack?from=discover-new', 'https://whitereaperaskingride.bandcamp.com/album/asking-for-a-ride?from=discover-new', 'https://collageeffect.bandcamp.com/album/emerald-network?from=discover-new', 'https://foxteethnj.bandcamp.com/album/through-the-blue?from=discover-new'] ['https://finitysounds.bandcamp.com/album/kreme?from=discover-new', 'https://mylittlerobotfriend.bandcamp.com/album/amen-break?from=discover-new', 'https://electrinityband.bandcamp.com/album/rise?from=discover-new', 'https://abyssal-void.bandcamp.com/album/ritualist?from=discover-new', 'https://plataformarecs.bandcamp.com/album/v-a-david-lynch-experience?from=discover-new', 'https://hurricaneturtles.bandcamp.com/album/industrial-synth?from=discover-new', 'https://blackwashband.bandcamp.com/album/2?from=discover-new', 'https://worldwide-bitchin-records.bandcamp.com/album/wack?from=discover-new']
解决方案
问题根源
- 旧元素缓存未清理:Selenium的
find_elements会保留之前页面的元素引用,用titles[-8:]取最后8个元素时,可能还是旧页面的DOM元素,没有重新定位当前页面的新内容。 - 等待条件无效:固定XPATH定位太依赖页面结构,Bandcamp翻页后DOM可能微调,导致等待的元素还是旧页面的,没真正触发新页面加载完成的判断。
- 翻页逻辑缺陷:翻页后没有验证页面是否真的更新,直接抓取内容,且页码元素可能还是旧页面的集合,点击操作可能未生效。
修复步骤
- 补上缺失的
time模块导入,代码里用了time.sleep但未声明。 - 调整流程:先抓取当前页内容,再执行翻页操作,避免页码变量逻辑混乱。
- 每次翻页后重新定位元素,不要依赖之前的元素集合,直接抓取当前页的8个目标元素。
- 优化等待条件:通过URL变化、新内容加载状态验证页面是否更新,避免无效等待。
- 用JS点击页码元素,规避页面元素遮挡导致的点击失败问题。
修复后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time browser = webdriver.Chrome() all_links = [] current_page = 1 url = "https://bandcamp.com/?g=all&s=new&p=0&gn=0&f=digital&w=-1" browser.get(url) wait = WebDriverWait(browser, 20) # 抓取第一页内容 wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "item-title"))) titles = browser.find_elements(By.CLASS_NAME, "item-title") current_links = [title.get_attribute('href') for title in titles[:8]] all_links.extend(current_links) print(f"第{current_page}页链接: {current_links}") # 抓取第2到第5页 while current_page < 5: current_page += 1 # 等待页码元素加载,找到目标页码并点击 page_nums = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "item-page"))) for page_num in page_nums: if page_num.text.isnumeric() and int(page_num.text) == current_page: browser.execute_script("arguments[0].click();", page_num) break # 等待页面URL更新,确认翻页完成 wait.until(lambda driver: driver.current_url != url) url = browser.current_url # 等待新页面专辑元素加载完成 wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "item-title"))) # 重新定位当前页的专辑链接 titles = browser.find_elements(By.CLASS_NAME, "item-title") current_links = [title.get_attribute('href') for title in titles[:8]] all_links.extend(current_links) print(f"第{current_page}页链接: {current_links}") browser.quit() print("所有链接: ", all_links)
额外建议
- 避免使用固定XPATH,改用
By.CLASS_NAME或By.CSS_SELECTOR这类更稳定的定位方式。 - 优先用
WebDriverWait动态等待,替代固定时长的time.sleep,提升脚本效率和稳定性。 - 可以添加重复链接校验逻辑,避免意外重复抓取。
内容的提问来源于stack exchange,提问作者Jrob1765
相关产品推荐
相关产品推荐

