You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页属性随机化时,如何用Python爬取novel5s小说页面内容?

爬取伪元素中的小说段落文本方案

问题说明

爬取指定小说页面时,段落文本并非普通DOM节点文本,而是存储在每次刷新会随机化的标签属性对应的::before和::after伪元素中,常规Selenium+BeautifulSoup方法无法提取完整内容。

解决方案

浏览器伪元素不属于DOM树节点,无法通过常规元素定位方法直接获取,需借助JavaScript的getComputedStyle方法获取伪元素的content属性值。同时注意:页面伪元素内容可能依赖JavaScript生成,需确保浏览器启用JS。

完整代码实现

from bs4 import BeautifulSoup
from selenium import webdriver
import chromedriver_autoinstaller
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
import time

# 自动安装ChromeDriver
chromedriver_autoinstaller.install()

# 配置Chrome选项,**不要禁用JavaScript**
chrome_options = Options()
# 可根据需求添加无头模式等其他配置,比如:
# chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)
driver.maximize_window()

try:
    driver.get("https://novel5s.com/bye-my-irresistible-love-by-goreous-novel5-online-2138/148981.html")
    time.sleep(3)  # 等待页面完全加载

    # 获取所有段落元素
    paragraphs = driver.find_elements(By.CSS_SELECTOR, ".content-book p")
    
    full_content = []
    for p in paragraphs:
        # 执行JS获取::before伪元素的content
        before_content = driver.execute_script("""
            return window.getComputedStyle(arguments[0], '::before').getPropertyValue('content');
        """, p)
        # 执行JS获取::after伪元素的content
        after_content = driver.execute_script("""
            return window.getComputedStyle(arguments[0], '::after').getPropertyValue('content');
        """, p)
        
        # 去除content两端的引号(伪元素content通常带引号包裹)
        before_text = before_content.strip('"\'')
        after_text = after_content.strip('"\'')
        
        # 拼接完整段落文本
        paragraph_text = before_text + after_text
        if paragraph_text:
            full_content.append(paragraph_text)
    
    # 输出完整内容
    print("\n".join(full_content))

finally:
    # 关闭浏览器
    driver.quit()

关键说明

  • 启用JavaScript:原代码中禁用JS的配置会导致伪元素内容无法生成,必须移除该设置
  • 伪元素获取逻辑:通过getComputedStyle指定伪元素类型(::before/::after),提取其content属性
  • 内容处理:伪元素的content值默认带引号,需用strip('"\'')去除多余引号
  • 异常处理:用try...finally确保浏览器无论是否报错都会正常关闭

内容的提问来源于stack exchange,提问作者user20205716

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 07:25:14