You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Scrapy-Splash采集截图元数据并写入Item?

Got it, let's work through this problem together. You're using Scrapy-Splash to take screenshots and need to attach specific metadata (target site IP, page load UTC timestamp, screenshot UTC timestamp) to your Scrapy Items—since Scrapy's internal engine logs don't expose this data directly in spiders, here's a solid solution:

Step 1: Define Your Scrapy Item

First, create an Item class to hold all your screenshot data and metadata:

import scrapy

class ScreenshotMetadataItem(scrapy.Item):
    image_data = scrapy.Field()  # Binary PNG data
    target_ip = scrapy.Field()
    page_load_utc = scrapy.Field()
    screenshot_utc = scrapy.Field()
    url = scrapy.Field()

Step 2: Modify Splash Lua Script to Return Metadata

The key trick is to extend your Splash Lua script to capture the timestamps and target IP directly from Splash's context (this avoids discrepancies between local DNS resolution and what Splash actually connects to). Then we'll return this data alongside the screenshot:

import json
import base64
import scrapy
from scrapy_splash import SplashRequest
from your_project.items import ScreenshotMetadataItem

class ScreenshotSpider(scrapy.Spider):
    name = "screenshot_crawler"
    allowed_domains = ["your-target-domain.com"]
    start_urls = ["https://your-target-domain.com"]

    def start_requests(self):
        for url in self.start_urls:
            # Lua script to capture metadata and screenshot
            splash_script = """
            function main(splash, args)
                -- Capture target IP by making a HEAD request first
                local head_req = splash:request{
                    url=args.url,
                    method="HEAD"
                }
                local target_ip = head_req.remote_addr

                -- Load the full page
                splash:go(args.url)
                splash:wait(3)  -- Adjust wait time based on page load speed

                -- Capture UTC timestamp when page finishes loading
                local page_load_time = os.date("!%Y-%m-%dT%H:%M:%SZ")

                -- Prepare full viewport screenshot
                splash:set_viewport_full()
                -- Capture UTC timestamp right before taking screenshot
                local screenshot_time = os.date("!%Y-%m-%dT%H:%M:%SZ")
                local screenshot = splash:png()

                -- Return all data as JSON
                return {
                    target_ip = target_ip,
                    page_load_utc = page_load_time,
                    screenshot_utc = screenshot_time,
                    screenshot = screenshot
                }
            end
            """

            yield SplashRequest(
                url=url,
                callback=self.parse_screenshot,
                endpoint="execute",
                args={"lua_source": splash_script, "url": url},
                meta={"original_url": url}
            )

    def parse_screenshot(self, response):
        item = ScreenshotMetadataItem()
        # Parse Splash's JSON response
        splash_data = json.loads(response.body)

        # Populate metadata fields
        item["target_ip"] = splash_data["target_ip"]
        item["page_load_utc"] = splash_data["page_load_utc"]
        item["screenshot_utc"] = splash_data["screenshot_utc"]
        item["url"] = response.meta["original_url"]

        # Decode base64 screenshot to binary data
        item["image_data"] = base64.b64decode(splash_data["screenshot"])

        yield item

Step 3: Add a Pipeline to Save Screenshots (Optional)

If you want to save the PNG files to disk with metadata in filenames, add a custom Pipeline:

import os
import base64
from scrapy.pipelines.images import ImagesPipeline
from scrapy.exceptions import DropItem
from scrapy.http import Request

class ScreenshotSavePipeline(ImagesPipeline):
    def file_path(self, request, response=None, info=None, item=None):
        # Use IP and screenshot timestamp for unique filenames
        safe_timestamp = item["screenshot_utc"].replace(":", "-")
        return f"screenshots/{item['target_ip']}_{safe_timestamp}.png"

    def get_media_requests(self, item, info):
        # Wrap binary image data in a data URI request for the pipeline
        data_uri = f"data:image/png;base64,{base64.b64encode(item['image_data']).decode('utf-8')}"
        yield Request(data_uri, meta={"item": item})

    def item_completed(self, results, item, info):
        image_paths = [x["path"] for ok, x in results if ok]
        if not image_paths:
            raise DropItem("No screenshot found for item")
        item["saved_path"] = image_paths[0]
        return item

Then enable the pipeline in your settings.py:

ITEM_PIPELINES = {
    "your_project.pipelines.ScreenshotSavePipeline": 300,
}

IMAGES_STORE = "./"  # Saves screenshots to a 'screenshots' subfolder

Bonus: Add Scrapy Engine Timestamps (If Needed)

If you also want to include Scrapy's internal request/response timestamps, you can pull them from the response's meta dict:

from datetime import datetime

def parse_screenshot(self, response):
    # ... existing code ...

    # Add Scrapy's request/response UTC timestamps
    request_timestamp = response.meta.get("downloader/request_time")
    if request_timestamp:
        item["scrapy_request_utc"] = datetime.utcfromtimestamp(request_timestamp).strftime("%Y-%m-%dT%H:%M:%SZ")
    
    response_timestamp = response.meta.get("downloader/response_time")
    if response_timestamp:
        item["scrapy_response_utc"] = datetime.utcfromtimestamp(response_timestamp).strftime("%Y-%m-%dT%H:%M:%SZ")

    yield item

Key Notes

  • Make sure your Splash service is running and configured correctly in settings.py (check SPLASH_URL, DOWNLOADER_MIDDLEWARES, and SPIDER_MIDDLEWARES).
  • Adjust splash:wait() or use splash:wait_for_selector() if your target pages load dynamically (e.g., React/Vue apps).
  • Using Splash's remote_addr gives you the actual IP Splash connected to, which is more accurate than local DNS resolution (especially if Splash is running in a container/VPC with different DNS settings).

内容的提问来源于stack exchange,提问作者CLPatterson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:53:16