如何用Scrapy与Python 2.7爬取指定网页并批量下载图片
Got it, let's walk through building this Scrapy spider step by step. I've broken it down into clear, actionable parts so you can get up and running quickly.
First, open your terminal and create a new Scrapy project and spider:
# Create project scrapy startproject time_covers # Navigate into project directory cd time_covers # Generate a spider targeting the time.com domain scrapy genspider time_cover_spider content.time.com
Open items.py in your project folder and add a custom item to handle image data. This is required for Scrapy's built-in image download pipeline:
import scrapy class TimeCoverItem(scrapy.Item): # Field to store the image URL(s) image_urls = scrapy.Field() # Field where Scrapy will store metadata about downloaded images images = scrapy.Field()
Open spiders/time_cover_spider.py and replace the default code with this. It handles image extraction, JSON output, and pagination:
import scrapy from time_covers.items import TimeCoverItem class TimeCoverSpider(scrapy.Spider): name = 'time_cover_spider' # Start with the URL you provided start_urls = ['http://content.time.com/time/covers/0,16641,19230303,00.html'] def parse(self, response): # Extract the cover image URL (adjust the CSS selector if needed) cover_image_url = response.css('div.cover-image img::attr(src)').get() if cover_image_url: # Create an Item and populate the image URL field item = TimeCoverItem() # Wrap the URL in a list (required by the ImagesPipeline) item['image_urls'] = [cover_image_url] yield item # Handle pagination: Find the "Next" button URL next_page = response.css('a.next::attr(href)').get() if next_page: # Convert relative URL to absolute URL (in case it's not full path) next_page_url = response.urljoin(next_page) # Recursively call parse on the next page yield scrapy.Request(url=next_page_url, callback=self.parse)
Note: If the CSS selector for the cover image doesn't work, inspect the page to find the correct selector (right-click the image → Inspect Element).
Open settings.py and update these sections to enable image downloads and set up output:
# Set a user agent to mimic a browser (avoids basic anti-scraping blocks) USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' # Define where downloaded images will be stored IMAGES_STORE = './time_covers_images' # Enable the ImagesPipeline to handle image downloads ITEM_PIPELINES = { 'scrapy.pipelines.images.ImagesPipeline': 1, }
Execute this command in your terminal. It will:
- Download all cover images to the
time_covers_imagesfolder - Save all image URLs in JSON format to
covers.json
scrapy crawl time_cover_spider -o covers.json
- If the next page isn't being detected: Double-check the CSS selector for the "Next" button (inspect the button element to find the correct class/attribute)
- If images aren't downloading: Ensure the
IMAGES_STOREpath is valid, and that the image URLs extracted are full, working URLs - If you hit anti-scraping blocks: Add a
DOWNLOAD_DELAYinsettings.py(e.g.,DOWNLOAD_DELAY = 2) to slow down requests
内容的提问来源于stack exchange,提问作者Dhawal Shukal

