Web Scraping图片抓取问题:无法找到'rel'选择器及脚本执行报错
Fixing Your Swords Comic Scraper IndexError & Image Download Issues
Let's walk through what's tripping up your script and fix it—you're almost there!
What's Going Wrong
- Comic Image Selector Mismatch: Your script uses
#comic-imageto find the comic, but on inner pages (like/comic/CCCLXII/), the comic image is nested inside a different container (#comic imginstead of a direct#comic-imageelement). That's why you see "Could not find comic image" on the second page. - Flawed URL Handling: Your manual string manipulation for
comicUrlis error-prone and creates broken URLs (like the double slash inhttp://swordscomic.com//media/...). - Unsafe Prev Button Selection: You directly access
soup.select('a[id=navigation-previous]')[0]without checking if the list has elements. On some pages (or if the selector is wrong), this triggers thatIndexError. Also, the site doesn't use thatidfor the previous button—it usesrel="prev"instead.
Fixed Script
Here's the revised code with key changes explained:
#! python3 # swordscraper.py - Downloads all the swords comics. import requests, os, bs4 from urllib.parse import urljoin # For safe URL joining # Set up your save directory os.chdir(r'C:\Users\bromp\OneDrive\Desktop\Python') os.makedirs('swords', exist_ok=True) base_url = 'https://swordscomic.com/' current_url = base_url # Starting URL while True: print(f'Downloading page {current_url}...') # Download and parse the page res = requests.get(current_url) res.raise_for_status() # Fixed: you missed parentheses here! soup = bs4.BeautifulSoup(res.text, 'html.parser') # 1. Find the comic image (works for both home and inner pages) comic_elem = soup.select('#comic img') if comic_elem: comic_src = comic_elem[0].get('src') # Use urljoin to handle relative/absolute URLs correctly comic_url = urljoin(base_url, comic_src) print(f'Downloading image {comic_url}...') img_res = requests.get(comic_url) img_res.raise_for_status() # Save the image img_filename = os.path.basename(comic_url) with open(os.path.join('swords', img_filename), 'wb') as img_file: for chunk in img_res.iter_content(100000): img_file.write(chunk) else: print(f'Could not find comic image on {current_url}') # 2. Find the "Previous" link (using rel="prev" which is consistent across pages) prev_links = soup.select('a[rel="prev"]') if not prev_links: print('No more previous pages found. Done!') break # Update current_url with the safe joined URL current_url = urljoin(base_url, prev_links[0].get('href'))
Key Changes Explained
- Fixed Error Checking: You had
res.raise_for_status(missing parentheses), which meant it wasn't actually validating HTTP responses. Now it properly throws an error if a page download fails. - Universal Comic Selector:
#comic imgtargets the image inside the main comic container, working for both the homepage and inner comic pages. - Safe URL Handling:
urljoinautomatically handles relative paths (like/comic/CCCLXII/) and absolute URLs, so you don't need messy string replacements. - Robust Prev Button Check: We first verify if
prev_linkshas elements before accessing the index, eliminating theIndexError. The site usesrel="prev"for previous links—a standard, consistent attribute across all pages. - Cleaner File Handling: Using a
withstatement ensures the image file is properly closed without manualclose()calls.
This should let your script download all comics without getting stuck or throwing index errors.
内容的提问来源于stack exchange,提问作者Brompy
相关产品推荐
相关产品推荐

