如何从带缩放功能的Canvas中抓取完整高清图像?
Hey there, let's figure out how to grab that full high-resolution image from the Canvas-based DeepZoom viewer you're working with! I’ve run into this exact issue before with OpenSeadragon (the library powering that viewer), so here’s what’s going on and how to fix it.
The problem with your existing approach is simple: OpenSeadragon only renders the image tiles visible in your browser’s viewport to the Canvas. When you call toDataURL(), you’re only capturing that small, zoomed-in section—not the full high-resolution image.
Since the image is served as tiled DeepZoom assets, we can use Selenium to fetch the metadata file, then download all the high-res tiles and stitch them together. Here’s how to do it step by step:
Step 1: Fetch the DeepZoom Metadata URL (.dzi) with Selenium
First, we’ll use Selenium to execute JavaScript that grabs the viewer’s metadata URL. OpenSeadragon exposes a global instance we can access directly:
# Get the .dzi metadata URL using Selenium dzi_url = driver.execute_script(""" // Grab the OpenSeadragon viewer instance const viewer = OpenSeadragon.getInstance(); // Return the URL to the DeepZoom metadata file return viewer.source.dziUrl; """) print("Metadata URL:", dzi_url)
Step 2: Parse the Metadata to Get Image Details
Next, we’ll download and parse the .dzi XML file to get key details like total image dimensions, tile size, and the highest resolution level:
import requests import xml.etree.ElementTree as ET import math # Download the metadata XML response = requests.get(dzi_url) root = ET.fromstring(response.content) # Extract core image properties image_props = root.find('Image') tile_size = int(image_props.get('TileSize')) image_format = image_props.get('Format') size_props = image_props.find('Size') total_width = int(size_props.get('Width')) total_height = int(size_props.get('Height')) # Calculate the highest resolution level (max zoom) max_level = math.ceil(math.log(max(total_width, total_height) / tile_size, 2)) print(f"Full image size: {total_width}x{total_height}, Max resolution level: {max_level}")
Step 3: Download All High-Resolution Tiles
Now we’ll construct URLs for every tile in the highest resolution level and download them. We’ll use requests for efficiency, but you could also use Selenium if you need to stay within the browser context:
import os from PIL import Image # Create a temp folder to store tiles temp_tile_dir = "deepzoom_tiles" os.makedirs(temp_tile_dir, exist_ok=True) # Calculate number of tiles needed num_cols = math.ceil(total_width / tile_size) num_rows = math.ceil(total_height / tile_size) # Download each tile (add headers to mimic browser traffic if needed) for col in range(num_cols): for row in range(num_rows): # Build the tile URL from the metadata URL tile_url = dzi_url.replace(".dzi", f"_files/{max_level}/{col}_{row}.{image_format}") # Use Selenium's user-agent to avoid anti-crawl blocks tile_response = requests.get(tile_url, headers={"User-Agent": driver.execute_script("return navigator.userAgent;")}) # Save the tile to temp folder tile_path = os.path.join(temp_tile_dir, f"{col}_{row}.{image_format}") with open(tile_path, "wb") as f: f.write(tile_response.content) print(f"Downloaded tile {col}_{row}")
Step 4: Stitch Tiles Into the Full High-Resolution Image
Finally, we’ll use Pillow to combine all the tiles into one complete image:
# Create a blank image with the full dimensions full_image = Image.new("RGB", (total_width, total_height)) # Paste each tile into the correct position for col in range(num_cols): for row in range(num_rows): tile_path = os.path.join(temp_tile_dir, f"{col}_{row}.{image_format}") tile = Image.open(tile_path) # Adjust position for edge tiles that are smaller than full tile size x_pos = col * tile_size if col != num_cols - 1 else total_width - tile.width y_pos = row * tile_size if row != num_rows - 1 else total_height - tile.height full_image.paste(tile, (x_pos, y_pos)) # Save the final full-res image full_image.save("full_highres_image.png") print("Full high-resolution image saved as full_highres_image.png")
- Anti-Crawl Measures: If you hit 403 errors, add more browser-like headers (like
Referer) to your requests, matching what you see in your browser’s network tab. - Tile Overlap: If your metadata shows an
Overlapvalue greater than 0, crop each tile (except edge tiles) by the overlap amount before pasting to avoid duplicate edges. - Large Images: For extremely big images, process tiles in batches to avoid running out of memory.
内容的提问来源于stack exchange,提问作者AYehia0

