如何通过Scrapy-Splash采集截图元数据并写入Item?
Got it, let's work through this problem together. You're using Scrapy-Splash to take screenshots and need to attach specific metadata (target site IP, page load UTC timestamp, screenshot UTC timestamp) to your Scrapy Items—since Scrapy's internal engine logs don't expose this data directly in spiders, here's a solid solution:
Step 1: Define Your Scrapy Item
First, create an Item class to hold all your screenshot data and metadata:
import scrapy class ScreenshotMetadataItem(scrapy.Item): image_data = scrapy.Field() # Binary PNG data target_ip = scrapy.Field() page_load_utc = scrapy.Field() screenshot_utc = scrapy.Field() url = scrapy.Field()
Step 2: Modify Splash Lua Script to Return Metadata
The key trick is to extend your Splash Lua script to capture the timestamps and target IP directly from Splash's context (this avoids discrepancies between local DNS resolution and what Splash actually connects to). Then we'll return this data alongside the screenshot:
import json import base64 import scrapy from scrapy_splash import SplashRequest from your_project.items import ScreenshotMetadataItem class ScreenshotSpider(scrapy.Spider): name = "screenshot_crawler" allowed_domains = ["your-target-domain.com"] start_urls = ["https://your-target-domain.com"] def start_requests(self): for url in self.start_urls: # Lua script to capture metadata and screenshot splash_script = """ function main(splash, args) -- Capture target IP by making a HEAD request first local head_req = splash:request{ url=args.url, method="HEAD" } local target_ip = head_req.remote_addr -- Load the full page splash:go(args.url) splash:wait(3) -- Adjust wait time based on page load speed -- Capture UTC timestamp when page finishes loading local page_load_time = os.date("!%Y-%m-%dT%H:%M:%SZ") -- Prepare full viewport screenshot splash:set_viewport_full() -- Capture UTC timestamp right before taking screenshot local screenshot_time = os.date("!%Y-%m-%dT%H:%M:%SZ") local screenshot = splash:png() -- Return all data as JSON return { target_ip = target_ip, page_load_utc = page_load_time, screenshot_utc = screenshot_time, screenshot = screenshot } end """ yield SplashRequest( url=url, callback=self.parse_screenshot, endpoint="execute", args={"lua_source": splash_script, "url": url}, meta={"original_url": url} ) def parse_screenshot(self, response): item = ScreenshotMetadataItem() # Parse Splash's JSON response splash_data = json.loads(response.body) # Populate metadata fields item["target_ip"] = splash_data["target_ip"] item["page_load_utc"] = splash_data["page_load_utc"] item["screenshot_utc"] = splash_data["screenshot_utc"] item["url"] = response.meta["original_url"] # Decode base64 screenshot to binary data item["image_data"] = base64.b64decode(splash_data["screenshot"]) yield item
Step 3: Add a Pipeline to Save Screenshots (Optional)
If you want to save the PNG files to disk with metadata in filenames, add a custom Pipeline:
import os import base64 from scrapy.pipelines.images import ImagesPipeline from scrapy.exceptions import DropItem from scrapy.http import Request class ScreenshotSavePipeline(ImagesPipeline): def file_path(self, request, response=None, info=None, item=None): # Use IP and screenshot timestamp for unique filenames safe_timestamp = item["screenshot_utc"].replace(":", "-") return f"screenshots/{item['target_ip']}_{safe_timestamp}.png" def get_media_requests(self, item, info): # Wrap binary image data in a data URI request for the pipeline data_uri = f"data:image/png;base64,{base64.b64encode(item['image_data']).decode('utf-8')}" yield Request(data_uri, meta={"item": item}) def item_completed(self, results, item, info): image_paths = [x["path"] for ok, x in results if ok] if not image_paths: raise DropItem("No screenshot found for item") item["saved_path"] = image_paths[0] return item
Then enable the pipeline in your settings.py:
ITEM_PIPELINES = { "your_project.pipelines.ScreenshotSavePipeline": 300, } IMAGES_STORE = "./" # Saves screenshots to a 'screenshots' subfolder
Bonus: Add Scrapy Engine Timestamps (If Needed)
If you also want to include Scrapy's internal request/response timestamps, you can pull them from the response's meta dict:
from datetime import datetime def parse_screenshot(self, response): # ... existing code ... # Add Scrapy's request/response UTC timestamps request_timestamp = response.meta.get("downloader/request_time") if request_timestamp: item["scrapy_request_utc"] = datetime.utcfromtimestamp(request_timestamp).strftime("%Y-%m-%dT%H:%M:%SZ") response_timestamp = response.meta.get("downloader/response_time") if response_timestamp: item["scrapy_response_utc"] = datetime.utcfromtimestamp(response_timestamp).strftime("%Y-%m-%dT%H:%M:%SZ") yield item
Key Notes
- Make sure your Splash service is running and configured correctly in
settings.py(checkSPLASH_URL,DOWNLOADER_MIDDLEWARES, andSPIDER_MIDDLEWARES). - Adjust
splash:wait()or usesplash:wait_for_selector()if your target pages load dynamically (e.g., React/Vue apps). - Using Splash's
remote_addrgives you the actual IP Splash connected to, which is more accurate than local DNS resolution (especially if Splash is running in a container/VPC with different DNS settings).
内容的提问来源于stack exchange,提问作者CLPatterson

