You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python加载完整动态HTML页面并获取指定前缀链接?

Got it, let's tackle this problem step by step. The issue you're facing is super common with dynamically loaded pages—requests only grabs the initial HTML sent by the server, and the rest of the content gets loaded via JavaScript as you scroll. Here are two solid approaches to get all those links you need:


方法1:用Selenium模拟浏览器滚动加载完整页面

Since the content loads as you scroll, we can use Selenium to simulate a real browser session, scroll to the bottom repeatedly until no new content loads, then extract the full page source.

Step 1: Install dependencies

First, install Selenium and download the matching browser driver (e.g., ChromeDriver for Chrome):

pip install selenium

(Make sure the driver version matches your browser's version, and place it in a directory accessible by your Python script.)

Step 2: Code implementation

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import time

url = "your_my_url_here"

# Set up Chrome options (headless mode runs without a visible window)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")

# Initialize driver and load page
driver = webdriver.Chrome(options=chrome_options)
driver.get(url)

# Simulate scrolling to load all content
last_scroll_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # Scroll to the bottom of the page
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # Wait for new content to load (adjust sleep time based on page speed)
    time.sleep(3)
    # Check if scroll height changed (no new content = break loop)
    new_scroll_height = driver.execute_script("return document.body.scrollHeight")
    if new_scroll_height == last_scroll_height:
        break
    last_scroll_height = new_scroll_height

# Get full page source and close driver
full_page_source = driver.page_source
driver.quit()

# Parse and extract target links
soup = BeautifulSoup(full_page_source, "html.parser")
target_links = []
for link in soup.find_all("a", href=True):
    href = link["href"]
    # Match links starting with your target pattern
    if href.startswith(f"{url}/something/"):
        target_links.append(href)

print("Found target links:")
for link in target_links:
    print(link)

方法2:直接抓取后端API接口(更高效)

Most scroll-loaded pages fetch content via AJAX/Fetch requests to a backend API. Instead of simulating a browser, you can directly call these APIs to get the raw data (usually JSON), which is faster and uses fewer resources.

Step 1: Find the API endpoint

  1. Open your page in a browser, press F12 to open DevTools, and go to the Network tab.
  2. Filter requests by XHR/Fetch (this shows dynamic data requests).
  3. Scroll the page and look for requests that return the content you need. Check the response payload—you’ll likely see the links embedded in JSON.
  4. Note the API URL, request method (GET/POST), and any required parameters (like page, offset, or limit).

Step 2: Code implementation

import requests

your_my_url = "your_my_url_here"
api_url = "https://example.com/api/load-more-items"  # Replace with your found API endpoint
page = 1
target_links = []

while True:
    # Adjust parameters based on the API you found
    params = {
        "page": page,
        "limit": 20  # Common parameter for number of items per page
    }
    # Add headers if needed (e.g., User-Agent, Authorization)
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    
    response = requests.get(api_url, params=params, headers=headers)
    data = response.json()
    
    # Extract links from the JSON structure (adjust based on actual data)
    items = data.get("items", [])
    if not items:
        break  # No more data to load
    
    for item in items:
        link = item.get("url")  # Replace with the actual key for links in the JSON
        if link and link.startswith(f"{your_my_url}/something/"):
            target_links.append(link)
    
    page += 1

print("Found target links:")
for link in target_links:
    print(link)

Which method should you choose?

  • Use Selenium if the page has complex JavaScript rendering (e.g., anti-scraping measures, dynamic DOM changes) and you can’t easily find the API.
  • Use the API approach if you can identify the backend requests—it’s faster, lighter, and more reliable long-term.

内容的提问来源于stack exchange,提问作者solopiu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 20:02:54