如何使用Python脚本分析图像格式的各类图表并对比图表颜色的相似性与差异性
Comparing Chart Colors Across Images with Python
Great question! Comparing chart colors to find similarities and differences is totally doable with Python, using a mix of image processing and color analysis libraries. Here's a practical, step-by-step approach with a complete script you can adapt to your needs:
Required Libraries
First, install the packages we’ll use for image handling, clustering, and visualization:
pip install opencv-python numpy scikit-learn matplotlib
Approach Overview
We’ll break this down into 4 core steps:
- Load and preprocess your images
- Isolate each of the 5 charts in every image
- Extract the dominant color from each chart
- Compare colors across images to measure similarity and highlight differences
Complete Customizable Script
Here’s a fully commented script that you can tweak based on your image layout and color sensitivity needs:
import cv2 import numpy as np from sklearn.cluster import KMeans import matplotlib.pyplot as plt from typing import List, Tuple def load_image(image_path: str) -> np.ndarray: """Load an image and convert from OpenCV's default BGR to RGB.""" img = cv2.imread(image_path) return cv2.cvtColor(img, cv2.COLOR_BGR2RGB) def extract_chart_regions(image: np.ndarray, chart_coords: List[Tuple[int, int, int, int]]) -> List[np.ndarray]: """Crop the image to isolate each chart using (x, y, width, height) coordinates.""" charts = [] for (x, y, w, h) in chart_coords: chart = image[y:y+h, x:x+w] charts.append(chart) return charts def get_dominant_color(image: np.ndarray, num_colors: int = 1) -> Tuple[int, int, int]: """Extract the most frequent color from a chart using K-means clustering.""" # Reshape image into a 2D array of pixels pixels = image.reshape(-1, 3) # Fit K-means to find dominant color clusters kmeans = KMeans(n_clusters=num_colors, random_state=42) kmeans.fit(pixels) # Pick the cluster with the most pixels (dominant color) counts = np.bincount(kmeans.labels_) dominant_idx = np.argmax(counts) return tuple(map(int, kmeans.cluster_centers_[dominant_idx])) def color_similarity(color1: Tuple[int, int, int], color2: Tuple[int, int, int], use_hsv: bool = True) -> float: """Calculate similarity between two colors (lower value = more similar).""" if use_hsv: # Convert to HSV for human-perceptual color comparison c1 = cv2.cvtColor(np.uint8([[color1]]), cv2.COLOR_RGB2HSV)[0][0] c2 = cv2.cvtColor(np.uint8([[color2]]), cv2.COLOR_RGB2HSV)[0][0] # Weight hue more heavily (it's the primary "color" component) hue_diff = abs(c1[0] - c2[0]) hue_diff = min(hue_diff, 180 - hue_diff) # Handle circular hue range (0-179) sat_diff = abs(c1[1] - c2[1]) val_diff = abs(c1[2] - c2[2]) return (hue_diff * 2) + sat_diff + val_diff else: # Euclidean distance in RGB space (less perceptually accurate) return np.linalg.norm(np.array(color1) - np.array(color2)) def analyze_chart_colors(image_paths: List[str], chart_coords: List[Tuple[int, int, int, int]]) -> dict: """Analyze all images and return dominant colors for each chart.""" results = {} for img_path in image_paths: img = load_image(img_path) charts = extract_chart_regions(img, chart_coords) colors = [get_dominant_color(chart) for chart in charts] results[img_path] = colors return results def compare_results(results: dict, similarity_threshold: int = 20) -> None: """Generate text summary and visualizations of color matches/differences.""" image_names = list(results.keys()) all_colors = list(results.values()) # Print comparison summary print("Chart Color Comparison Summary:\n") for i in range(len(image_names)): for j in range(i+1, len(image_names)): img1, img2 = image_names[i], image_names[j] matches, differences = [], [] for chart_idx in range(5): sim_score = color_similarity(all_colors[i][chart_idx], all_colors[j][chart_idx]) if sim_score < similarity_threshold: matches.append(f"Chart {chart_idx+1} ({all_colors[i][chart_idx]} = {all_colors[j][chart_idx]})") else: differences.append(f"Chart {chart_idx+1} ({all_colors[i][chart_idx]} vs {all_colors[j][chart_idx]}, score: {sim_score:.1f})") print(f"Between {img1} and {img2}:") if matches: print(f" Matching charts: {', '.join(matches)}") if differences: print(f" Different charts: {', '.join(differences)}") print() # Visualize chart colors for easy side-by-side comparison fig, axes = plt.subplots(len(results), 1, figsize=(10, len(results)*2)) for idx, (img_name, colors) in enumerate(results.items()): ax = axes[idx] if len(results) > 1 else axes ax.set_title(f"Chart Colors: {img_name}") ax.imshow([colors]) ax.set_yticks([]) ax.set_xticks(range(5)) ax.set_xticklabels([f"Chart {i+1}" for i in range(5)]) plt.tight_layout() plt.show() if __name__ == "__main__": # -------------------------- # Customize these parameters # -------------------------- # Paths to your image files IMAGE_PATHS = ["image1.jpg", "image2.jpg", "image3.jpg"] # Coordinates for each chart (x, y, width, height) # Use an image editor (GIMP/Paint) to get these values for your images CHART_COORDS = [ (50, 50, 100, 100), # Chart 1 (200, 50, 100, 100), # Chart 2 (350, 50, 100, 100), # Chart 3 (500, 50, 100, 100), # Chart 4 (650, 50, 100, 100) # Chart 5 ] # Similarity threshold (lower = stricter color match) SIMILARITY_THRESHOLD = 20 # Run the analysis color_results = analyze_chart_colors(IMAGE_PATHS, CHART_COORDS) compare_results(color_results, SIMILARITY_THRESHOLD)
Key Customizations for Your Use Case
- Chart Coordinates: Use an image editor to find the exact (x, y, width, height) of each chart in your images. Ensure these coordinates are consistent across all images.
- Dominant Color Handling: If your charts have borders or minor color variations, set
num_colors=2inget_dominant_colorand add logic to exclude the border color. - Similarity Threshold: Adjust
SIMILARITY_THRESHOLDbased on how strict you want matches to be. Lower values mean colors must be nearly identical. - Color Space: The script uses HSV for comparison (more human-friendly), but you can switch to RGB by setting
use_hsv=Falseincolor_similarity.
How It Works
- Image Loading: OpenCV loads images and converts them to RGB (since OpenCV defaults to BGR).
- Chart Isolation: We crop each image to isolate individual charts using your specified coordinates.
- Dominant Color Extraction: K-means clustering identifies the most frequent color in each chart, even if there’s slight variation.
- Similarity Calculation: Weighted HSV distance ensures we prioritize hue (the actual "color") over saturation/value.
- Comparison Output: The script prints a text summary of matches/differences and generates a visual grid of all chart colors for quick comparison.
If you need help adapting this to your specific image layout or color edge cases, feel free to ask!
内容的提问来源于stack exchange,提问作者dgs
相关产品推荐
相关产品推荐

