You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法抓取指定class的div解析歌曲标题,爬虫新手求助

问题解决:无法爬取网页中的歌曲信息div元素

核心问题分析

你的代码存在两个关键问题导致返回空列表:

  • BeautifulSoup查找class时,字符串末尾多了一个多余的单引号:"single-post-oembed-youtube-wrapper'" 应改为 "single-post-oembed-youtube-wrapper"
  • 未等待页面完全加载就获取源码,Selenium打开页面后直接提取page_source可能导致目标元素还未渲染完成

修正后的代码

import json
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.common.exceptions import TimeoutException
import pprint

url = 'https://ultimateclassicrock.com/best-rock-songs-2018/'

# 初始化浏览器(推荐用webdriver-manager自动管理驱动,避免路径问题)
browser = webdriver.Chrome()
browser.get(url)

try:
    # 显式等待目标元素加载,最长等待10秒
    WebDriverWait(browser, 10).until(
        EC.presence_of_all_elements_located((By.CLASS_NAME, 'single-post-oembed-youtube-wrapper'))
    )
    
    # 提取渲染后的页面源码
    soup = BeautifulSoup(browser.page_source, 'html.parser')
    # 修正class名的单引号错误
    divs = soup.find_all("div", {"class": "single-post-oembed-youtube-wrapper"})
    
    # 提取每个div对应的歌曲名和歌手信息(实际信息在相邻的figcaption标签中)
    song_info = []
    for div in divs:
        figcaption = div.find_next_sibling('figcaption')
        if figcaption:
            song_info.append(figcaption.get_text(strip=True))
    
    pprint.pprint(song_info)
except TimeoutException:
    print("页面加载超时,未找到目标元素")
finally:
    browser.close()

额外说明

  • 目标div本身是youtube视频容器,歌曲名和歌手信息实际在它后面的<figcaption>标签里,需通过find_next_sibling提取
  • 使用WebDriverWait显式等待元素加载,比直接获取源码更可靠,避免动态渲染导致的元素缺失
  • 推荐使用webdriver-manager自动管理ChromeDriver,替换硬编码的驱动路径,避免环境差异问题:
    from webdriver_manager.chrome import ChromeDriverManager
    browser = webdriver.Chrome(ChromeDriverManager().install())
    

内容的提问来源于stack exchange,提问作者Radamand

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 14:35:22