如何等待网站加载完所有HTML内容后获取响应与动态加载元素?
Hey there! I totally get your frustration with dynamically loaded content—those elements that pop up a few seconds after the initial page loads can be tricky to grab when you’re just fetching raw HTML. Let’s break down how to solve this so you can get that tab-match-head-2-head content reliably.
The Core Issue
When you use a basic HTTP request (like with Python’s requests library), you only get the initial HTML the server sends. The head-to-head tab content you’re after is loaded later by JavaScript running in the browser—so it doesn’t exist in that initial response. You need a way to simulate a real browser that waits for all that JS to execute and content to render, just like how window.onload works in the browser itself.
Solution 1: Use Selenium (Headless Browser)
Selenium mimics a real browser, so it can wait for dynamic elements to load before you grab their content. Here’s how to implement it:
First, install the required packages:
pip install selenium
You’ll also need a webdriver (like ChromeDriver) that matches your browser version—most modern browsers have built-in support now, so you might not need to download it separately.
Then write your code to wait for the target element:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC url = "https://www.scoreboard.com/game/berankis-ricardas-king-kevin-2018/WC4oWAqE/#h2h;all" # Launch Chrome in headless mode (no visible window) options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) try: driver.get(url) # Wait up to 10 seconds for the element to appear in the DOM target_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "tab-match-head-2-head")) ) # Grab the inner HTML you need inner_html = target_element.get_attribute("innerHTML") print(inner_html) finally: # Always close the browser when done driver.quit()
The WebDriverWait ensures we don’t try to access the element until it’s actually loaded—no more incomplete content!
Solution 2: Playwright (Modern, Reliable Alternative)
Playwright is a newer tool that’s designed specifically for handling dynamic web content. It’s often faster and more intuitive than Selenium. Here’s how to use it:
First, install Playwright and the required browser:
pip install playwright playwright install chrome
Then your code:
from playwright.sync_api import sync_playwright url = "https://www.scoreboard.com/game/berankis-ricardas-king-kevin-2018/WC4oWAqE/#h2h;all" with sync_playwright() as p: # Launch headless Chrome browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto(url) # Wait for the element to be visible (not just present in the DOM) page.wait_for_selector("#tab-match-head-2-head", state="visible") # Get the inner HTML inner_html = page.locator("#tab-match-head-2-head").inner_html() print(inner_html) browser.close()
Playwright’s built-in waits handle most edge cases automatically, so you don’t have to write as much boilerplate code.
Solution 3: Check for Hidden APIs (Advanced, More Efficient)
Sometimes websites load dynamic content via behind-the-scenes API calls. If you can find that API endpoint, you can fetch the data directly without using a browser at all—this is way faster. Here’s how to find it:
- Open the page in Chrome/Firefox and hit F12 to open DevTools.
- Go to the Network tab and filter by "XHR" or "Fetch".
- Refresh the page and look for requests that return the head-to-head data (check the response body).
- Copy that request’s URL and any required headers, then use
requeststo call it directly.
This skips the entire browser rendering step, which is ideal if you only need the raw data.
内容的提问来源于stack exchange,提问作者Olexandr Isakov

