使用Python+OpenCV2实现视频子集匹配时,读视频遇Killed 9崩溃求助
Hey there, let's get straight to it: that Killed 9 error you're seeing is almost always the OS stepping in to terminate your process because it's using too much memory (an out-of-memory, or OOM, condition). While vidcap.read() itself isn't the direct cause, how you're storing those frames is definitely the problem here.
Why Your Current Approach Is Crashing
Storing every single frame from multiple videos in dictionaries or lists is a surefire way to eat up RAM fast. Let's do a quick reality check: a 720p frame (1280x720, 3 color channels) is roughly 1.3MB per frame. A 40-second video at 30fps has 1200 frames—that's ~1.5GB just for one video. Throw in another video (even a 2-second clip adds 60 frames) and you're quickly pushing past your system's available memory. The OS's OOM killer then sends a SIGKILL (signal 9) to shut down your script before it crashes the whole system.
Better, Memory-Efficient Solutions
You don't need to store every frame to check if VideoA is a subset of VideoB. Here are some smarter approaches:
- Key Frame Hashes (Most Efficient)
Instead of processing every frame, extract "key frames" (frames that show significant content changes, or just skip every N frames) and compute unique hashes for them. Then you just need to check if all of VideoA's key frame hashes exist in VideoB's set of hashes.
Example code snippet:
import cv2 import imagehash from PIL import Image def get_key_frame_hashes(video_path, skip_interval=5): frame_hashes = set() cap = cv2.VideoCapture(video_path) frame_count = 0 while cap.isOpened(): success, frame = cap.read() if not success: break # Only process every Nth frame to save resources if frame_count % skip_interval == 0: # Convert OpenCV's BGR frame to PIL's RGB format pil_frame = Image.fromarray(cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)) # Compute a difference hash (dHash) for the frame frame_hash = imagehash.dhash(pil_frame) frame_hashes.add(frame_hash) frame_count += 1 cap.release() return frame_hashes # Check subset status video_a_hashes = get_key_frame_hashes("videoA.mp4") video_b_hashes = get_key_frame_hashes("videoB.mp4") is_subset = video_a_hashes.issubset(video_b_hashes) print(f"VideoA is a subset of VideoB: {is_subset}")
Stream Frames On-the-Fly
Skip storing all of VideoB's frames entirely. Instead, stream VideoB frame by frame, and for each frame in VideoA, scan through VideoB until you find a matching frame (using hash similarity or pixel comparison). Once you find a match, continue checking the next VideoA frame right after that position in VideoB. This way, you only keep a few frames in memory at a time.Optimize If You Must Store Frames
If you absolutely need to hold frames in memory, compress them first (e.g., usecv2.imencode('.jpg', frame)to store JPEG bytes instead of raw numpy arrays) or resize frames to a smaller resolution (like 640x360) to cut memory usage drastically. Also, use generators to yield frames one at a time instead of collecting them all in a list.
Quick Best Practices
- Always call
cap.release()after you're done with aVideoCaptureobject to free up system resources. - Use tools like
top(Linux/macOS) or Task Manager (Windows) to monitor memory usage while your script runs—this will confirm if OOM is the issue. - For high-res videos, resize frames early in processing to reduce the data size you're working with.
内容的提问来源于stack exchange,提问作者donttellunclesam

