You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取动态渲染的Microsoft CVE-2018-8176页面内容?

Got it, since this page uses dynamic content rendering (loads data via JavaScript after the initial page load), regular requests won't work because it only fetches the static HTML source. Here are two reliable Python approaches to grab the description content:

使用Selenium

Selenium simulates a real browser, so it can execute JavaScript and wait for dynamic content to load.

  1. Install dependencies:

    pip install selenium
    

    You also need to download a browser driver matching your browser version (e.g., ChromeDriver for Google Chrome, GeckoDriver for Firefox).

  2. Sample code:

    from selenium import webdriver
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # Initialize Chrome browser (replace with Firefox/Edge if needed)
    driver = webdriver.Chrome()
    try:
        # Navigate to the target page
        driver.get("https://portal.msrc.microsoft.com/en-US/security-guidance/advisory/CVE-2018-8176")
        
        # Wait explicitly for the description section to load (max 10 seconds)
        description_element = WebDriverWait(driver, 10).until(
            EC.visibility_of_element_located((By.ID, "description"))
        )
        
        # Extract the text content of the description
        description_content = description_element.text
        print(description_content)
    finally:
        # Always close the browser to avoid resource leaks
        driver.quit()
    
  3. Tips:

    • Use headless mode for background execution: add options.add_argument("--headless=new") when initializing the driver.
    • Match your driver version exactly to your browser version to avoid compatibility issues.
使用Playwright

Playwright is a modern alternative to Selenium, with better built-in support for dynamic content and no need to manually download drivers.

  1. Install dependencies:

    pip install playwright
    playwright install  # This installs Chrome, Firefox, and WebKit browsers
    
  2. Sample code:

    from playwright.sync_api import sync_playwright
    
    with sync_playwright() as p:
        # Launch Chromium browser (set headless=True for background mode)
        browser = p.chromium.launch(headless=False)
        page = browser.new_page()
        
        # Load the target page
        page.goto("https://portal.msrc.microsoft.com/en-US/security-guidance/advisory/CVE-2018-8176")
        
        # Wait for the description element to be available and get its text
        description_content = page.wait_for_selector("#description").text_content()
        print(description_content)
        
        # Clean up
        browser.close()
    
  3. Tips:

    • Use async_playwright instead of sync_playwright for asynchronous tasks (great for bulk scraping).
    • Playwright automatically handles most waiting scenarios, so you rarely need manual delays.

内容的提问来源于stack exchange,提问作者elrond2194

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:21:28