使用Requests+BeautifulSoup无法获取页面img标签,如何高效替代Selenium?
Hey there! Let's break down why your current approach isn't working and share a way more efficient alternative than Selenium that matches the speed of BeautifulSoup.
Why Requests + BeautifulSoup Isn't Finding the Images
The issue here is that MangaDex loads chapter images dynamically with JavaScript. When you use requests.get(), you're only fetching the initial static HTML of the page—those img tags you see in your browser's dev tools don't exist in that initial response. They're added to the page later after the browser runs the site's JS, which fetches image data from MangaDex's API. That's why your soup doesn't pick them up, even though the status code is 200.
Efficient Alternative: Use MangaDex's Public API
Instead of scraping the rendered page, you can directly call MangaDex's official public API to get the chapter's image list. This is way faster than Selenium (no browser overhead) and just as efficient as your original Requests/BeautifulSoup setup. Here's how to do it:
Step-by-Step Implementation
- Extract the chapter ID from your URL (in your case, it's
435396fromhttps://mangadex.org/chapter/435396/2) - Call the chapter details API endpoint to get image metadata
- Construct the full image URLs using the returned data
Here's the updated code:
import requests # Your target chapter ID (extracted from the URL) chapter_id = "435396" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.76 Safari/537.36' } # 1. Fetch chapter details from MangaDex API api_url = f"https://api.mangadex.org/chapter/{chapter_id}" response = requests.get(api_url, headers=headers) response.raise_for_status() # Raise error if request fails chapter_data = response.json() # 2. Extract image metadata from the API response image_hash = chapter_data['data']['attributes']['hash'] image_files = chapter_data['data']['attributes']['data'] # 3. Build full image URLs image_urls = [f"https://uploads.mangadex.org/data/{image_hash}/{file}" for file in image_files] # Print all valid image URLs print("Chapter Image URLs:") for url in image_urls: print(url)
Key Advantages of This Method
- No HTML parsing needed: We get structured JSON data directly, so there's no dependency on BeautifulSoup for this task
- Speed & efficiency: Just a couple of lightweight API calls—far faster than spinning up a browser with Selenium
- Reliability: APIs are designed for programmatic access, so they're less likely to break if MangaDex updates their page layout
内容的提问来源于stack exchange,提问作者Doc

