You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium爬取网页时如何按指定分界元素对列表元素分组输出

问题原因

全局匹配所有class为_3PpPJ OrSDI的span元素,会将所有和弦、段落标记节点拉平为一维列表,完全丢失DOM原有的行级、段落级层级结构,无法保留原生排版。

实现方案

不需要单独匹配每个段落节点再查询同级后续元素,直接定位和弦谱根容器,逐行遍历容器内的行级元素即可还原排版:

  • 每行内的和弦span按页面顺序拼接,保留行内和弦排列
  • 自动识别方括号包裹的段落标记(如[Verse 1]、[Chorus])作为段落分界
  • 保留原页面的空行、行顺序,输出结构和网页显示完全一致
修正后代码
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
import undetected_chromedriver as uc
import time
from selenium.common.exceptions import NoSuchElementException

options = Options() 
options.add_argument("start-maximized")
options.add_argument('--no-sandbox') 
driver = uc.Chrome(options=options)

target_url = '目标吉他谱页面地址'
driver.get(target_url)
driver.implicitly_wait(30)

# 处理可能弹出的协议/广告弹窗
try:
    confirm_btn = driver.find_element(By.XPATH, '//button[contains(text(), "Accept") or contains(text(), "同意")]')
    confirm_btn.click()
    time.sleep(1)
except NoSuchElementException:
    pass

# 定位谱面根容器
tab_root = driver.find_element(By.XPATH, '//section[contains(@class, "js-tab-content")]')
# 遍历所有行元素
tab_lines = tab_root.find_elements(By.XPATH, './/div[contains(@class, "_3yg3h")]')
for line in tab_lines:
    span_list = line.find_elements(By.XPATH, './/span[@class="_3PpPJ OrSDI"]')
    line_content = ' '.join([s.text for s in span_list if s.text.strip()])
    print(line_content)

driver.quit()
注意事项
  • 如果页面后续改版导致行元素class_3yg3h失效,可在浏览器开发者工具中选中任意一行和弦,复制其对应父级行元素的class属性替换即可
  • 该写法自动适配所有段落类型,无需为Verse、Chorus、Bridge等不同段落单独编写匹配规则
  • 输出会自动保留原页面的空行间距、和弦顺序,无需额外排版处理

内容的提问来源于stack exchange,提问作者harley-

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 08:27:18