You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Scrapy从嵌套URL中爬取图片(附代码示例)

Solution for Handling Nested Image URLs in Scrapy Media Pipeline

Got it, let's tackle this nested image URL handling in Scrapy. The default ImagesPipeline works great for top-level image_urls, but since your images are nested inside the results array of each item, we need to build a custom pipeline to traverse those sub-items and handle the image downloads properly.

Step 1: Fix the image_urls Format in Your Spider

The ImagesPipeline expects image_urls to be an iterable (list/tuple), not a single string. Update your Spider's parse method to wrap the extracted URL in a list (and handle cases where no image is found):

# Replace this line
data["image_urls"] = caritem.css("div.view-auction a img::attr(src)").extract_first()
# With this
img_url = caritem.css("div.view-auction a img::attr(src)").extract_first()
data["image_urls"] = [img_url] if img_url else []

This ensures we pass a valid format to the pipeline, even if no image exists for a car item.

Step 2: Create a Custom Nested Images Pipeline

Add this custom pipeline class to your pipelines.py file. It inherits from ImagesPipeline and overrides key methods to process nested entries:

from scrapy.pipelines.images import ImagesPipeline
from scrapy.exceptions import DropItem
from scrapy.http import Request

class NestedImagesPipeline(ImagesPipeline):

    def get_media_requests(self, item, info):
        # Traverse each nested result in the main item
        for result in item.get('results', []):
            # Yield a request for each image URL in the nested result
            for image_url in result.get('image_urls', []):
                # Pass the nested result via meta to map downloads back later
                yield Request(image_url, meta={'result': result})

    def item_completed(self, results, item, info):
        # Map downloaded image paths back to their corresponding nested result
        for result in item.get('results', []):
            image_paths = []
            # Match each download result to the correct nested entry
            for ok, download_data in results:
                if ok and download_data['url'] in result.get('image_urls', []):
                    image_paths.append(download_data['path'])
            # Add the local image paths to the nested result (name this field whatever you prefer)
            result['images'] = image_paths
        return item

    def process_item(self, item, info):
        # Skip items that don't have the nested results field
        if 'results' not in item:
            raise DropItem("Item missing required 'results' field")
        return super().process_item(item, info)

Step 3: Update Settings to Use the Custom Pipeline

In settings.py, replace the default ImagesPipeline with your custom one (adjust the path if you saved the pipeline elsewhere):

ITEM_PIPELINES = {
    # Use our custom nested image pipeline instead of the default
    'auctions_results.pipelines.NestedImagesPipeline': 1,
    # Keep any other existing pipelines you have here
}

# Ensure your image storage directory is set
IMAGES_STORE = './downloaded_images'  # Or your preferred local path

Step 4: Verify the Output

After running your spider, your JSON output will now include an images field in each nested results entry, containing the local path to the downloaded image:

[
    {
        "gm_url": "example-url",
        "results": [
            {
                "marque": "ferrari",
                "auction_house": "auction",
                "auction_country": "japan",
                "auction_date": "2019",
                "image_urls": ["example-img-url"],
                "images": ["full/abc12345.jpg"]
            },
            ...
        ]
    }
]

Key Notes

  • Meta Tracking: We use meta={'result': result} in get_media_requests to link each image request to its parent nested result, so we can correctly map download paths back later.
  • Error Resilience: The pipeline skips invalid items and handles empty image URL lists gracefully.
  • Thumbnails: If you want generated thumbnails, add the IMAGES_THUMBS setting to your settings.py (e.g., IMAGES_THUMBS = {'small': (50, 50), 'large': (200, 200)}).

内容的提问来源于stack exchange,提问作者lf_celine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:28:38