You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy与Python 2.7爬取指定网页并批量下载图片

Got it, let's walk through building this Scrapy spider step by step. I've broken it down into clear, actionable parts so you can get up and running quickly.

1. Set Up the Scrapy Project

First, open your terminal and create a new Scrapy project and spider:

# Create project
scrapy startproject time_covers

# Navigate into project directory
cd time_covers

# Generate a spider targeting the time.com domain
scrapy genspider time_cover_spider content.time.com
2. Define the Item Structure

Open items.py in your project folder and add a custom item to handle image data. This is required for Scrapy's built-in image download pipeline:

import scrapy

class TimeCoverItem(scrapy.Item):
    # Field to store the image URL(s)
    image_urls = scrapy.Field()
    # Field where Scrapy will store metadata about downloaded images
    images = scrapy.Field()
3. Write the Spider Logic

Open spiders/time_cover_spider.py and replace the default code with this. It handles image extraction, JSON output, and pagination:

import scrapy
from time_covers.items import TimeCoverItem

class TimeCoverSpider(scrapy.Spider):
    name = 'time_cover_spider'
    # Start with the URL you provided
    start_urls = ['http://content.time.com/time/covers/0,16641,19230303,00.html']

    def parse(self, response):
        # Extract the cover image URL (adjust the CSS selector if needed)
        cover_image_url = response.css('div.cover-image img::attr(src)').get()
        
        if cover_image_url:
            # Create an Item and populate the image URL field
            item = TimeCoverItem()
            # Wrap the URL in a list (required by the ImagesPipeline)
            item['image_urls'] = [cover_image_url]
            yield item
        
        # Handle pagination: Find the "Next" button URL
        next_page = response.css('a.next::attr(href)').get()
        if next_page:
            # Convert relative URL to absolute URL (in case it's not full path)
            next_page_url = response.urljoin(next_page)
            # Recursively call parse on the next page
            yield scrapy.Request(url=next_page_url, callback=self.parse)

Note: If the CSS selector for the cover image doesn't work, inspect the page to find the correct selector (right-click the image → Inspect Element).

4. Configure Image Download & Settings

Open settings.py and update these sections to enable image downloads and set up output:

# Set a user agent to mimic a browser (avoids basic anti-scraping blocks)
USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'

# Define where downloaded images will be stored
IMAGES_STORE = './time_covers_images'

# Enable the ImagesPipeline to handle image downloads
ITEM_PIPELINES = {
    'scrapy.pipelines.images.ImagesPipeline': 1,
}
5. Run the Spider & Get Results

Execute this command in your terminal. It will:

  • Download all cover images to the time_covers_images folder
  • Save all image URLs in JSON format to covers.json
scrapy crawl time_cover_spider -o covers.json
Troubleshooting Tips
  • If the next page isn't being detected: Double-check the CSS selector for the "Next" button (inspect the button element to find the correct class/attribute)
  • If images aren't downloading: Ensure the IMAGES_STORE path is valid, and that the image URLs extracted are full, working URLs
  • If you hit anti-scraping blocks: Add a DOWNLOAD_DELAY in settings.py (e.g., DOWNLOAD_DELAY = 2) to slow down requests

内容的提问来源于stack exchange,提问作者Dhawal Shukal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:08:07