You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取时,为何dt标签文本无法输出?

问题原因及解决方案

为什么提取不到dt标签的文本?

核心原因有两个:

  • <template>标签的特殊性:浏览器默认不会渲染<template>内部的内容,尽管Selenium能抓取到该标签的HTML代码,但BeautifulSoup对这类未被渲染的模板节点解析时,容易出现文本识别异常。
  • 解析器兼容性问题:你使用的lxml解析器对<template>内带注释的文本处理存在偏差,导致无法正确识别标签内的文本节点。

可行的解决方案

方案1:用Selenium直接获取template内部HTML再解析

绕过BeautifulSoup对整个页面中template节点的解析问题,先通过Selenium提取template的内部HTML,再单独解析:

import time
import os
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By  # 新增导入

if __name__ == '__main__':
    os.environ['WDM_LOG'] = '0' 
    options = Options()
    options.add_argument("start-maximized")
    options.add_experimental_option("prefs", {"profile.default_content_setting_values.notifications": 1})    
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option('excludeSwitches', ['enable-logging'])
    options.add_experimental_option('useAutomationExtension', False)
    options.add_argument('--disable-blink-features=AutomationControlled') 
    srv=Service()
    driver = webdriver.Chrome(service=srv, options=options)        

    wLink = "https://www.medimops.de/agatha-christie-agatha-christie-ein-schritt-ins-leere-why-didn-t-they-ask-evans-der-komplette-vierteiler-mit-starbesetzung-blu-ray-blu-ray-M0B0BW28MKKR.html"
    driver.get(wLink)       
    time.sleep(3) 

    # 直接用Selenium定位template,获取内部HTML
    worker = driver.find_element(By.CLASS_NAME, "product-attributes__table")
    template = worker.find_element(By.TAG_NAME, "template")
    template_html = template.get_attribute('innerHTML')
    
    # 解析template内部的HTML
    soup = BeautifulSoup(template_html, 'lxml')
    wDT = soup.find("dt", {"class": "product-attributes__definition"})
    print(wDT.text.strip().replace(':', ''))  # 输出EAN / ISBN
    
    driver.quit()

方案2:更换BeautifulSoup的解析器

把lxml换成Python内置的html.parser,该解析器对模板节点的文本识别更稳定:

# 修改BeautifulSoup初始化的代码行
soup = BeautifulSoup(driver.page_source, 'html.parser')

方案3:手动遍历dt节点的子内容

直接提取dt标签下的所有文本节点,跳过注释干扰:

# 找到dt标签后,遍历其内容节点
wDT = worker.find("dt")
text_content = ''.join([node.strip() for node in wDT.contents if isinstance(node, str)]).replace(':', '').strip()
print(text_content)  # 输出EAN / ISBN

内容的提问来源于stack exchange,提问作者Rapid1898

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 23:09:50