You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法获取网页完整源码:Python爬虫请求库爬取问题求助

Solution for Missing Dynamic HTML in Web Scraping with Python

Hey there! I totally get your frustration—when you can see the content in your browser but can’t scrape it with requests or urllib, it almost always boils down to dynamic content loading. Let me break this down and show you how to fix it.

Why This Happens

The Udemy page you’re targeting loads its course data using JavaScript. When you use requests/urllib, you’re only fetching the initial static HTML that the server sends—none of the JavaScript has run yet to render the actual course cards or data-rich tags. Your browser, on the other hand, executes all the JS and builds the full DOM you see in DevTools.

Fix 1: Use Selenium to Simulate a Browser

Selenium launches a real (or headless) browser, lets it render the page fully, and then gives you access to the complete HTML. Here’s how to adjust your code:

First, install Selenium and a browser driver (like ChromeDriver):

pip install selenium

Make sure you have ChromeDriver downloaded and in your PATH (or specify its path in the code).

Then update your scraping code:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup as BS

url = 'https://www.udemy.com/topic/financial-analysis/?lang=en'
user_agent = 'your-actual-user-agent' # Replace with your real UA from DevTools

# Set up headless Chrome (so no browser window pops up)
chrome_options = Options()
chrome_options.add_argument(f'user-agent={user_agent}')
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--disable-gpu')

driver = webdriver.Chrome(options=chrome_options)
driver.get(url)

# Wait a few seconds to let all content load (or use explicit waits for better reliability)
driver.implicitly_wait(5)

# Get the fully rendered HTML
html = driver.page_source
soup = BS(html, 'html.parser')

# Now you can scrape the content you see in DevTools!
# Example: Find course titles
course_titles = soup.find_all('div', class_='course-card--course-title--vVEjC')
for title in course_titles:
    print(title.get_text(strip=True))

driver.quit()

Fix 2: Directly Call Udemy’s API (Faster & More Efficient)

Instead of scraping HTML, you can fetch the data directly from Udemy’s internal API. This is faster because you avoid rendering the whole page. Here’s how to find and use it:

  1. Open DevTools in your browser (F12) → Go to the Network tab → Filter by XHR
  2. Refresh the Udemy page, and look for requests that return course data (they usually have discovery-units or courses in the URL)
  3. Copy the request URL, headers, and parameters

Here’s a simplified example of calling the API with requests:

import requests

url = 'https://www.udemy.com/api-2.0/discovery-units/all_courses/?p=1&page_size=60&subcategory_id=&instructional_level=&lang=en&price=&duration=&closed_captions=&subtopic_id=&topic_id=165&source_page=topic_page&locale=en_US&currency=usd&navigation_locale=en_US&skip_price=true&sos=pc&fl=topic'
headers = {
    'User-Agent': 'your-actual-user-agent', # Use your real UA
    'Accept': 'application/json, text/plain, */*'
}

response = requests.get(url, headers=headers)
data = response.json()

# Extract course data from the JSON
for course in data['unit']['items']:
    print(f"Course Title: {course['title']}")
    print(f"URL: https://www.udemy.com{course['url']}")
    print(f"Rating: {course['rating']}")
    print("---")

Note: API endpoints and parameters might change over time, so you’ll need to recheck the Network tab if this stops working.

Quick Tip

Always use your actual browser’s User-Agent string (you can find it in DevTools under Network → Request Headers) to avoid getting blocked by Udemy’s anti-scraping measures.

内容的提问来源于stack exchange,提问作者Muhammad Nadeem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:28:19