使用BeautifulSoup爬取时,为何dt标签文本无法输出?
问题原因及解决方案
为什么提取不到dt标签的文本?
核心原因有两个:
<template>标签的特殊性:浏览器默认不会渲染<template>内部的内容,尽管Selenium能抓取到该标签的HTML代码,但BeautifulSoup对这类未被渲染的模板节点解析时,容易出现文本识别异常。- 解析器兼容性问题:你使用的
lxml解析器对<template>内带注释的文本处理存在偏差,导致无法正确识别标签内的文本节点。
可行的解决方案
方案1:用Selenium直接获取template内部HTML再解析
绕过BeautifulSoup对整个页面中template节点的解析问题,先通过Selenium提取template的内部HTML,再单独解析:
import time import os from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By # 新增导入 if __name__ == '__main__': os.environ['WDM_LOG'] = '0' options = Options() options.add_argument("start-maximized") options.add_experimental_option("prefs", {"profile.default_content_setting_values.notifications": 1}) options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('excludeSwitches', ['enable-logging']) options.add_experimental_option('useAutomationExtension', False) options.add_argument('--disable-blink-features=AutomationControlled') srv=Service() driver = webdriver.Chrome(service=srv, options=options) wLink = "https://www.medimops.de/agatha-christie-agatha-christie-ein-schritt-ins-leere-why-didn-t-they-ask-evans-der-komplette-vierteiler-mit-starbesetzung-blu-ray-blu-ray-M0B0BW28MKKR.html" driver.get(wLink) time.sleep(3) # 直接用Selenium定位template,获取内部HTML worker = driver.find_element(By.CLASS_NAME, "product-attributes__table") template = worker.find_element(By.TAG_NAME, "template") template_html = template.get_attribute('innerHTML') # 解析template内部的HTML soup = BeautifulSoup(template_html, 'lxml') wDT = soup.find("dt", {"class": "product-attributes__definition"}) print(wDT.text.strip().replace(':', '')) # 输出EAN / ISBN driver.quit()
方案2:更换BeautifulSoup的解析器
把lxml换成Python内置的html.parser,该解析器对模板节点的文本识别更稳定:
# 修改BeautifulSoup初始化的代码行 soup = BeautifulSoup(driver.page_source, 'html.parser')
方案3:手动遍历dt节点的子内容
直接提取dt标签下的所有文本节点,跳过注释干扰:
# 找到dt标签后,遍历其内容节点 wDT = worker.find("dt") text_content = ''.join([node.strip() for node in wDT.contents if isinstance(node, str)]).replace(':', '').strip() print(text_content) # 输出EAN / ISBN
内容的提问来源于stack exchange,提问作者Rapid1898
相关产品推荐
相关产品推荐

