You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取Bandcamp发现卡片链接时HTML不更新的问题

问题

我写了一个简单的网页爬虫脚本,想抓取Bandcamp平台发现卡片里的所有超链接,但遇到了重复抓取的问题:脚本能正确获取第1页的8个专辑链接,但第2-4页重复抓取第2页的内容,后续也有类似重复情况。虽然Selenium控制的浏览器页面已经更新(URL也变了),但Selenium无法识别HTML变化,导致重复获取相同内容。我试过延长等待时间、用WebDriverWait等待不同元素加载,都没用。

我的代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.select import Select
browser = webdriver.Chrome()

all_links = []
page = 1
url = "https://bandcamp.com/?g=all&s=new&p=0&gn=0&f=digital&w=-1"
browser.get(url)
while page < 6:
    page += 1

    # wait until discover cards are loaded
    test = WebDriverWait(browser, 20).until(EC.element_to_be_clickable(
        (By.XPATH, '//*[@id="discover"]/div[9]/div[2]/div/div[1]/div/table/tbody/tr[1]/td[1]/a/div')))

    # scrape hyperlinks for each of the 8 albums shown
    titles = browser.find_elements(By.CLASS_NAME, "item-title")
    links = [title.get_attribute('href') for title in titles[-8:]]
    all_links = all_links + links
    print(links)

    # pagination - click through the page buttons as the links are scraped
    page_nums = browser.find_elements(By.CLASS_NAME, 'item-page')
    for page_num in page_nums:
        if page_num.text.isnumeric():
            if int(page_num.text) == page:
                page_num.click()
                time.sleep(20) # I've tried multiple long wait times as well as WebDriverWaits on different elements to see if the HTML will update, but I haven't seen a positive effect
                break

输出的重复链接示例:

['https://neubauten.bandcamp.com/album/stimmen-reste-musterhaus-7?from=discover-new', 'https://cirka1.bandcamp.com/album/time?from=discover-new', 'https://futuramusicsound.bandcamp.com/album/yoga-meditation?from=discover-new', 'https://deathsoundbatrecordings.bandcamp.com/album/real-mushrooms-dsbep092?from=discover-new', 'https://riacurley.bandcamp.com/album/take-me-album?from=discover-new', 'https://terracuna.bandcamp.com/album/el-origen-del-viento?from=discover-new', 'https://hyper-music.bandcamp.com/album/hypermusic-vol-4?from=discover-new', 'https://defisis1.bandcamp.com/album/priceless?from=discover-new']
['https://jarnosalo.bandcamp.com/album/here-lies-ancient-blob?from=discover-new', 'https://andreneitzel.bandcamp.com/album/allegasi-gold-2?from=discover-new', 'https://moonraccoon.bandcamp.com/album/prequels?from=discover-new', 'https://lolivone.bandcamp.com/album/live-at-the-berklee-performance-center?from=discover-new', 'https://nilswrasse.bandcamp.com/album/a-calling-from-the-desert-to-the-sea-original-motion-picture-soundtrack?from=discover-new', 'https://whitereaperaskingride.bandcamp.com/album/asking-for-a-ride?from=discover-new', 'https://collageeffect.bandcamp.com/album/emerald-network?from=discover-new', 'https://foxteethnj.bandcamp.com/album/through-the-blue?from=discover-new']
['https://jarnosalo.bandcamp.com/album/here-lies-ancient-blob?from=discover-new', 'https://andreneitzel.bandcamp.com/album/allegasi-gold-2?from=discover-new', 'https://moonraccoon.bandcamp.com/album/prequels?from=discover-new', 'https://lolivone.bandcamp.com/album/live-at-the-berklee-performance-center?from=discover-new', 'https://nilswrasse.bandcamp.com/album/a-calling-from-the-desert-to-the-sea-original-motion-picture-soundtrack?from=discover-new', 'https://whitereaperaskingride.bandcamp.com/album/asking-for-a-ride?from=discover-new', 'https://collageeffect.bandcamp.com/album/emerald-network?from=discover-new', 'https://foxteethnj.bandcamp.com/album/through-the-blue?from=discover-new']
['https://jarnosalo.bandcamp.com/album/here-lies-ancient-blob?from=discover-new', 'https://andreneitzel.bandcamp.com/album/allegasi-gold-2?from=discover-new', 'https://moonraccoon.bandcamp.com/album/prequels?from=discover-new', 'https://lolivone.bandcamp.com/album/live-at-the-berklee-performance-center?from=discover-new', 'https://nilswrasse.bandcamp.com/album/a-calling-from-the-desert-to-the-sea-original-motion-picture-soundtrack?from=discover-new', 'https://whitereaperaskingride.bandcamp.com/album/asking-for-a-ride?from=discover-new', 'https://collageeffect.bandcamp.com/album/emerald-network?from=discover-new', 'https://foxteethnj.bandcamp.com/album/through-the-blue?from=discover-new']
['https://finitysounds.bandcamp.com/album/kreme?from=discover-new', 'https://mylittlerobotfriend.bandcamp.com/album/amen-break?from=discover-new', 'https://electrinityband.bandcamp.com/album/rise?from=discover-new', 'https://abyssal-void.bandcamp.com/album/ritualist?from=discover-new', 'https://plataformarecs.bandcamp.com/album/v-a-david-lynch-experience?from=discover-new', 'https://hurricaneturtles.bandcamp.com/album/industrial-synth?from=discover-new', 'https://blackwashband.bandcamp.com/album/2?from=discover-new', 'https://worldwide-bitchin-records.bandcamp.com/album/wack?from=discover-new']
解决方案

问题根源

  1. 旧元素缓存未清理:Selenium的find_elements会保留之前页面的元素引用,用titles[-8:]取最后8个元素时,可能还是旧页面的DOM元素,没有重新定位当前页面的新内容。
  2. 等待条件无效:固定XPATH定位太依赖页面结构,Bandcamp翻页后DOM可能微调,导致等待的元素还是旧页面的,没真正触发新页面加载完成的判断。
  3. 翻页逻辑缺陷:翻页后没有验证页面是否真的更新,直接抓取内容,且页码元素可能还是旧页面的集合,点击操作可能未生效。

修复步骤

  1. 补上缺失的time模块导入,代码里用了time.sleep但未声明。
  2. 调整流程:先抓取当前页内容,再执行翻页操作,避免页码变量逻辑混乱。
  3. 每次翻页后重新定位元素,不要依赖之前的元素集合,直接抓取当前页的8个目标元素。
  4. 优化等待条件:通过URL变化、新内容加载状态验证页面是否更新,避免无效等待。
  5. 用JS点击页码元素,规避页面元素遮挡导致的点击失败问题。

修复后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

browser = webdriver.Chrome()
all_links = []
current_page = 1
url = "https://bandcamp.com/?g=all&s=new&p=0&gn=0&f=digital&w=-1"
browser.get(url)

wait = WebDriverWait(browser, 20)
# 抓取第一页内容
wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "item-title")))
titles = browser.find_elements(By.CLASS_NAME, "item-title")
current_links = [title.get_attribute('href') for title in titles[:8]]
all_links.extend(current_links)
print(f"第{current_page}页链接: {current_links}")

# 抓取第2到第5页
while current_page < 5:
    current_page += 1
    # 等待页码元素加载,找到目标页码并点击
    page_nums = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "item-page")))
    for page_num in page_nums:
        if page_num.text.isnumeric() and int(page_num.text) == current_page:
            browser.execute_script("arguments[0].click();", page_num)
            break
    
    # 等待页面URL更新,确认翻页完成
    wait.until(lambda driver: driver.current_url != url)
    url = browser.current_url
    
    # 等待新页面专辑元素加载完成
    wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "item-title")))
    # 重新定位当前页的专辑链接
    titles = browser.find_elements(By.CLASS_NAME, "item-title")
    current_links = [title.get_attribute('href') for title in titles[:8]]
    all_links.extend(current_links)
    print(f"第{current_page}页链接: {current_links}")

browser.quit()
print("所有链接: ", all_links)

额外建议

  • 避免使用固定XPATH,改用By.CLASS_NAME或By.CSS_SELECTOR这类更稳定的定位方式。
  • 优先用WebDriverWait动态等待,替代固定时长的time.sleep,提升脚本效率和稳定性。
  • 可以添加重复链接校验逻辑,避免意外重复抓取。

内容的提问来源于stack exchange,提问作者Jrob1765

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 11:55:41