如何用Python加载完整动态HTML页面并获取指定前缀链接?
Got it, let's tackle this problem step by step. The issue you're facing is super common with dynamically loaded pages—requests only grabs the initial HTML sent by the server, and the rest of the content gets loaded via JavaScript as you scroll. Here are two solid approaches to get all those links you need:
Since the content loads as you scroll, we can use Selenium to simulate a real browser session, scroll to the bottom repeatedly until no new content loads, then extract the full page source.
Step 1: Install dependencies
First, install Selenium and download the matching browser driver (e.g., ChromeDriver for Chrome):
pip install selenium
(Make sure the driver version matches your browser's version, and place it in a directory accessible by your Python script.)
Step 2: Code implementation
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time url = "your_my_url_here" # Set up Chrome options (headless mode runs without a visible window) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") # Initialize driver and load page driver = webdriver.Chrome(options=chrome_options) driver.get(url) # Simulate scrolling to load all content last_scroll_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to the bottom of the page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for new content to load (adjust sleep time based on page speed) time.sleep(3) # Check if scroll height changed (no new content = break loop) new_scroll_height = driver.execute_script("return document.body.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height # Get full page source and close driver full_page_source = driver.page_source driver.quit() # Parse and extract target links soup = BeautifulSoup(full_page_source, "html.parser") target_links = [] for link in soup.find_all("a", href=True): href = link["href"] # Match links starting with your target pattern if href.startswith(f"{url}/something/"): target_links.append(href) print("Found target links:") for link in target_links: print(link)
Most scroll-loaded pages fetch content via AJAX/Fetch requests to a backend API. Instead of simulating a browser, you can directly call these APIs to get the raw data (usually JSON), which is faster and uses fewer resources.
Step 1: Find the API endpoint
- Open your page in a browser, press F12 to open DevTools, and go to the Network tab.
- Filter requests by XHR/Fetch (this shows dynamic data requests).
- Scroll the page and look for requests that return the content you need. Check the response payload—you’ll likely see the links embedded in JSON.
- Note the API URL, request method (GET/POST), and any required parameters (like
page,offset, orlimit).
Step 2: Code implementation
import requests your_my_url = "your_my_url_here" api_url = "https://example.com/api/load-more-items" # Replace with your found API endpoint page = 1 target_links = [] while True: # Adjust parameters based on the API you found params = { "page": page, "limit": 20 # Common parameter for number of items per page } # Add headers if needed (e.g., User-Agent, Authorization) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(api_url, params=params, headers=headers) data = response.json() # Extract links from the JSON structure (adjust based on actual data) items = data.get("items", []) if not items: break # No more data to load for item in items: link = item.get("url") # Replace with the actual key for links in the JSON if link and link.startswith(f"{your_my_url}/something/"): target_links.append(link) page += 1 print("Found target links:") for link in target_links: print(link)
Which method should you choose?
- Use Selenium if the page has complex JavaScript rendering (e.g., anti-scraping measures, dynamic DOM changes) and you can’t easily find the API.
- Use the API approach if you can identify the backend requests—it’s faster, lighter, and more reliable long-term.
内容的提问来源于stack exchange,提问作者solopiu

