如何用Python抓取动态渲染的Microsoft CVE-2018-8176页面内容?
Got it, since this page uses dynamic content rendering (loads data via JavaScript after the initial page load), regular requests won't work because it only fetches the static HTML source. Here are two reliable Python approaches to grab the description content:
Selenium simulates a real browser, so it can execute JavaScript and wait for dynamic content to load.
Install dependencies:
pip install seleniumYou also need to download a browser driver matching your browser version (e.g., ChromeDriver for Google Chrome, GeckoDriver for Firefox).
Sample code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Initialize Chrome browser (replace with Firefox/Edge if needed) driver = webdriver.Chrome() try: # Navigate to the target page driver.get("https://portal.msrc.microsoft.com/en-US/security-guidance/advisory/CVE-2018-8176") # Wait explicitly for the description section to load (max 10 seconds) description_element = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.ID, "description")) ) # Extract the text content of the description description_content = description_element.text print(description_content) finally: # Always close the browser to avoid resource leaks driver.quit()Tips:
- Use headless mode for background execution: add
options.add_argument("--headless=new")when initializing the driver. - Match your driver version exactly to your browser version to avoid compatibility issues.
- Use headless mode for background execution: add
Playwright is a modern alternative to Selenium, with better built-in support for dynamic content and no need to manually download drivers.
Install dependencies:
pip install playwright playwright install # This installs Chrome, Firefox, and WebKit browsersSample code:
from playwright.sync_api import sync_playwright with sync_playwright() as p: # Launch Chromium browser (set headless=True for background mode) browser = p.chromium.launch(headless=False) page = browser.new_page() # Load the target page page.goto("https://portal.msrc.microsoft.com/en-US/security-guidance/advisory/CVE-2018-8176") # Wait for the description element to be available and get its text description_content = page.wait_for_selector("#description").text_content() print(description_content) # Clean up browser.close()Tips:
- Use
async_playwrightinstead ofsync_playwrightfor asynchronous tasks (great for bulk scraping). - Playwright automatically handles most waiting scenarios, so you rarely need manual delays.
- Use
内容的提问来源于stack exchange,提问作者elrond2194

