如何基于Scrapy从嵌套URL中爬取图片(附代码示例)
Got it, let's tackle this nested image URL handling in Scrapy. The default ImagesPipeline works great for top-level image_urls, but since your images are nested inside the results array of each item, we need to build a custom pipeline to traverse those sub-items and handle the image downloads properly.
Step 1: Fix the image_urls Format in Your Spider
The ImagesPipeline expects image_urls to be an iterable (list/tuple), not a single string. Update your Spider's parse method to wrap the extracted URL in a list (and handle cases where no image is found):
# Replace this line data["image_urls"] = caritem.css("div.view-auction a img::attr(src)").extract_first() # With this img_url = caritem.css("div.view-auction a img::attr(src)").extract_first() data["image_urls"] = [img_url] if img_url else []
This ensures we pass a valid format to the pipeline, even if no image exists for a car item.
Step 2: Create a Custom Nested Images Pipeline
Add this custom pipeline class to your pipelines.py file. It inherits from ImagesPipeline and overrides key methods to process nested entries:
from scrapy.pipelines.images import ImagesPipeline from scrapy.exceptions import DropItem from scrapy.http import Request class NestedImagesPipeline(ImagesPipeline): def get_media_requests(self, item, info): # Traverse each nested result in the main item for result in item.get('results', []): # Yield a request for each image URL in the nested result for image_url in result.get('image_urls', []): # Pass the nested result via meta to map downloads back later yield Request(image_url, meta={'result': result}) def item_completed(self, results, item, info): # Map downloaded image paths back to their corresponding nested result for result in item.get('results', []): image_paths = [] # Match each download result to the correct nested entry for ok, download_data in results: if ok and download_data['url'] in result.get('image_urls', []): image_paths.append(download_data['path']) # Add the local image paths to the nested result (name this field whatever you prefer) result['images'] = image_paths return item def process_item(self, item, info): # Skip items that don't have the nested results field if 'results' not in item: raise DropItem("Item missing required 'results' field") return super().process_item(item, info)
Step 3: Update Settings to Use the Custom Pipeline
In settings.py, replace the default ImagesPipeline with your custom one (adjust the path if you saved the pipeline elsewhere):
ITEM_PIPELINES = { # Use our custom nested image pipeline instead of the default 'auctions_results.pipelines.NestedImagesPipeline': 1, # Keep any other existing pipelines you have here } # Ensure your image storage directory is set IMAGES_STORE = './downloaded_images' # Or your preferred local path
Step 4: Verify the Output
After running your spider, your JSON output will now include an images field in each nested results entry, containing the local path to the downloaded image:
[ { "gm_url": "example-url", "results": [ { "marque": "ferrari", "auction_house": "auction", "auction_country": "japan", "auction_date": "2019", "image_urls": ["example-img-url"], "images": ["full/abc12345.jpg"] }, ... ] } ]
Key Notes
- Meta Tracking: We use
meta={'result': result}inget_media_requeststo link each image request to its parent nested result, so we can correctly map download paths back later. - Error Resilience: The pipeline skips invalid items and handles empty image URL lists gracefully.
- Thumbnails: If you want generated thumbnails, add the
IMAGES_THUMBSsetting to yoursettings.py(e.g.,IMAGES_THUMBS = {'small': (50, 50), 'large': (200, 200)}).
内容的提问来源于stack exchange,提问作者lf_celine

