You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests+BeautifulSoup无法获取页面img标签,如何高效替代Selenium?

Hey there! Let's break down why your current approach isn't working and share a way more efficient alternative than Selenium that matches the speed of BeautifulSoup.

Why Requests + BeautifulSoup Isn't Finding the Images

The issue here is that MangaDex loads chapter images dynamically with JavaScript. When you use requests.get(), you're only fetching the initial static HTML of the page—those img tags you see in your browser's dev tools don't exist in that initial response. They're added to the page later after the browser runs the site's JS, which fetches image data from MangaDex's API. That's why your soup doesn't pick them up, even though the status code is 200.

Efficient Alternative: Use MangaDex's Public API

Instead of scraping the rendered page, you can directly call MangaDex's official public API to get the chapter's image list. This is way faster than Selenium (no browser overhead) and just as efficient as your original Requests/BeautifulSoup setup. Here's how to do it:

Step-by-Step Implementation

  1. Extract the chapter ID from your URL (in your case, it's 435396 from https://mangadex.org/chapter/435396/2)
  2. Call the chapter details API endpoint to get image metadata
  3. Construct the full image URLs using the returned data

Here's the updated code:

import requests

# Your target chapter ID (extracted from the URL)
chapter_id = "435396"
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/56.0.2924.76 Safari/537.36'
}

# 1. Fetch chapter details from MangaDex API
api_url = f"https://api.mangadex.org/chapter/{chapter_id}"
response = requests.get(api_url, headers=headers)
response.raise_for_status()  # Raise error if request fails
chapter_data = response.json()

# 2. Extract image metadata from the API response
image_hash = chapter_data['data']['attributes']['hash']
image_files = chapter_data['data']['attributes']['data']

# 3. Build full image URLs
image_urls = [f"https://uploads.mangadex.org/data/{image_hash}/{file}" for file in image_files]

# Print all valid image URLs
print("Chapter Image URLs:")
for url in image_urls:
    print(url)

Key Advantages of This Method

  • No HTML parsing needed: We get structured JSON data directly, so there's no dependency on BeautifulSoup for this task
  • Speed & efficiency: Just a couple of lightweight API calls—far faster than spinning up a browser with Selenium
  • Reliability: APIs are designed for programmatic access, so they're less likely to break if MangaDex updates their page layout

内容的提问来源于stack exchange,提问作者Doc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:48:22