You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping图片抓取问题:无法找到'rel'选择器及脚本执行报错

Fixing Your Swords Comic Scraper IndexError & Image Download Issues

Let's walk through what's tripping up your script and fix it—you're almost there!

What's Going Wrong

  • Comic Image Selector Mismatch: Your script uses #comic-image to find the comic, but on inner pages (like /comic/CCCLXII/), the comic image is nested inside a different container (#comic img instead of a direct #comic-image element). That's why you see "Could not find comic image" on the second page.
  • Flawed URL Handling: Your manual string manipulation for comicUrl is error-prone and creates broken URLs (like the double slash in http://swordscomic.com//media/...).
  • Unsafe Prev Button Selection: You directly access soup.select('a[id=navigation-previous]')[0] without checking if the list has elements. On some pages (or if the selector is wrong), this triggers that IndexError. Also, the site doesn't use that id for the previous button—it uses rel="prev" instead.

Fixed Script

Here's the revised code with key changes explained:

#! python3
# swordscraper.py - Downloads all the swords comics.
import requests, os, bs4
from urllib.parse import urljoin  # For safe URL joining

# Set up your save directory
os.chdir(r'C:\Users\bromp\OneDrive\Desktop\Python')
os.makedirs('swords', exist_ok=True)

base_url = 'https://swordscomic.com/'
current_url = base_url  # Starting URL

while True:
    print(f'Downloading page {current_url}...')
    # Download and parse the page
    res = requests.get(current_url)
    res.raise_for_status()  # Fixed: you missed parentheses here!
    soup = bs4.BeautifulSoup(res.text, 'html.parser')

    # 1. Find the comic image (works for both home and inner pages)
    comic_elem = soup.select('#comic img')
    if comic_elem:
        comic_src = comic_elem[0].get('src')
        # Use urljoin to handle relative/absolute URLs correctly
        comic_url = urljoin(base_url, comic_src)
        
        print(f'Downloading image {comic_url}...')
        img_res = requests.get(comic_url)
        img_res.raise_for_status()
        
        # Save the image
        img_filename = os.path.basename(comic_url)
        with open(os.path.join('swords', img_filename), 'wb') as img_file:
            for chunk in img_res.iter_content(100000):
                img_file.write(chunk)
    else:
        print(f'Could not find comic image on {current_url}')

    # 2. Find the "Previous" link (using rel="prev" which is consistent across pages)
    prev_links = soup.select('a[rel="prev"]')
    if not prev_links:
        print('No more previous pages found. Done!')
        break
    
    # Update current_url with the safe joined URL
    current_url = urljoin(base_url, prev_links[0].get('href'))

Key Changes Explained

  • Fixed Error Checking: You had res.raise_for_status (missing parentheses), which meant it wasn't actually validating HTTP responses. Now it properly throws an error if a page download fails.
  • Universal Comic Selector: #comic img targets the image inside the main comic container, working for both the homepage and inner comic pages.
  • Safe URL Handling: urljoin automatically handles relative paths (like /comic/CCCLXII/) and absolute URLs, so you don't need messy string replacements.
  • Robust Prev Button Check: We first verify if prev_links has elements before accessing the index, eliminating the IndexError. The site uses rel="prev" for previous links—a standard, consistent attribute across all pages.
  • Cleaner File Handling: Using a with statement ensures the image file is properly closed without manual close() calls.

This should let your script download all comics without getting stuck or throwing index errors.

内容的提问来源于stack exchange,提问作者Brompy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 10:33:12