无法获取网页完整源码:Python爬虫请求库爬取问题求助
Hey there! I totally get your frustration—when you can see the content in your browser but can’t scrape it with requests or urllib, it almost always boils down to dynamic content loading. Let me break this down and show you how to fix it.
Why This Happens
The Udemy page you’re targeting loads its course data using JavaScript. When you use requests/urllib, you’re only fetching the initial static HTML that the server sends—none of the JavaScript has run yet to render the actual course cards or data-rich tags. Your browser, on the other hand, executes all the JS and builds the full DOM you see in DevTools.
Fix 1: Use Selenium to Simulate a Browser
Selenium launches a real (or headless) browser, lets it render the page fully, and then gives you access to the complete HTML. Here’s how to adjust your code:
First, install Selenium and a browser driver (like ChromeDriver):
pip install selenium
Make sure you have ChromeDriver downloaded and in your PATH (or specify its path in the code).
Then update your scraping code:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup as BS url = 'https://www.udemy.com/topic/financial-analysis/?lang=en' user_agent = 'your-actual-user-agent' # Replace with your real UA from DevTools # Set up headless Chrome (so no browser window pops up) chrome_options = Options() chrome_options.add_argument(f'user-agent={user_agent}') chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') driver = webdriver.Chrome(options=chrome_options) driver.get(url) # Wait a few seconds to let all content load (or use explicit waits for better reliability) driver.implicitly_wait(5) # Get the fully rendered HTML html = driver.page_source soup = BS(html, 'html.parser') # Now you can scrape the content you see in DevTools! # Example: Find course titles course_titles = soup.find_all('div', class_='course-card--course-title--vVEjC') for title in course_titles: print(title.get_text(strip=True)) driver.quit()
Fix 2: Directly Call Udemy’s API (Faster & More Efficient)
Instead of scraping HTML, you can fetch the data directly from Udemy’s internal API. This is faster because you avoid rendering the whole page. Here’s how to find and use it:
- Open DevTools in your browser (F12) → Go to the Network tab → Filter by XHR
- Refresh the Udemy page, and look for requests that return course data (they usually have
discovery-unitsorcoursesin the URL) - Copy the request URL, headers, and parameters
Here’s a simplified example of calling the API with requests:
import requests url = 'https://www.udemy.com/api-2.0/discovery-units/all_courses/?p=1&page_size=60&subcategory_id=&instructional_level=&lang=en&price=&duration=&closed_captions=&subtopic_id=&topic_id=165&source_page=topic_page&locale=en_US¤cy=usd&navigation_locale=en_US&skip_price=true&sos=pc&fl=topic' headers = { 'User-Agent': 'your-actual-user-agent', # Use your real UA 'Accept': 'application/json, text/plain, */*' } response = requests.get(url, headers=headers) data = response.json() # Extract course data from the JSON for course in data['unit']['items']: print(f"Course Title: {course['title']}") print(f"URL: https://www.udemy.com{course['url']}") print(f"Rating: {course['rating']}") print("---")
Note: API endpoints and parameters might change over time, so you’ll need to recheck the Network tab if this stops working.
Quick Tip
Always use your actual browser’s User-Agent string (you can find it in DevTools under Network → Request Headers) to avoid getting blocked by Udemy’s anti-scraping measures.
内容的提问来源于stack exchange,提问作者Muhammad Nadeem

