Python中ThreadPoolExecutor处理图片仅完成一半的原因咨询
Problem Description
I'm following Corey Schafer's multiprocessing/multithreading tutorials on YouTube. I'm working on an example where I need to add a blur effect to 15 images and save them to a processed subdirectory in the current folder.
The code works perfectly with ProcessPoolExecutor—all 15 images are processed and saved correctly. But when I replace ProcessPoolExecutor with ThreadPoolExecutor, only about half of the images get processed and saved. All images are in the same directory as my Python script.
Here's my code:
import time import concurrent.futures from PIL import Image, ImageFilter img_names = [ 'photo-1516117172878-fd2c41f4a759.jpg', 'photo-1532009324734-20a7a5813719.jpg', 'photo-1524429656589-6633a470097c.jpg', 'photo-1530224264768-7ff8c1789d79.jpg', 'photo-1564135624576-c5c88640f235.jpg', 'photo-1541698444083-023c97d3f4b6.jpg', 'photo-1522364723953-452d3431c267.jpg', 'photo-1513938709626-033611b8cc03.jpg', 'photo-1507143550189-fed454f93097.jpg', 'photo-1493976040374-85c8e12f0c0e.jpg', 'photo-1504198453319-5ce911bafcde.jpg', 'photo-1530122037265-a5f1f91d3b99.jpg', 'photo-1516972810927-80185027ca84.jpg', 'photo-1550439062-609e1531270e.jpg', 'photo-1549692520-acc6669e2f0c.jpg' ] def process_image(img_name): size = (1200,1200) img = Image.open(img_name) img = img.filter(ImageFilter.GaussianBlur(15)) img.thumbnail(size) img.save(f'processed/{img_name}') print(f"{img_name} was processed") def main(): t1 = time.perf_counter() with concurrent.futures.ThreadPoolExecutor() as executor: executor.map(process_image,img_names) t2 = time.perf_counter() print(f"finished in {t2-t1} seconds") if __name__ == "__main__": main()
Answer
Great question! This issue boils down to two key factors related to how threads work in CPython and Pillow's thread safety:
Pillow isn't fully thread-safe for certain operations
Pillow's underlying image processing operations (like applying Gaussian blur or saving files) rely on C extensions that aren't always thread-safe. When multiple threads try to execute these operations simultaneously, you can run into silent resource conflicts, memory errors, or unhandled exceptions that cause a thread to terminate without completing its task. Since these failures happen quietly, it looks like only half your images are being processed.You're not handling exceptions from
executor.map()
Theexecutor.map()method wraps any exceptions thrown by your worker function in the iterator it returns. If you don't iterate over this iterator (which your current code doesn't do), those exceptions get swallowed entirely. You never see an error message, and the failed tasks just disappear without a trace.
Why does ProcessPoolExecutor work? Because each process has its own isolated Python interpreter and memory space. There's no shared state between processes, so Pillow's internal resources don't get corrupted by concurrent access. All tasks run independently and complete successfully.
Fixes to Try
Option 1: Add exception handling to your worker function
This will let you see exactly which images are failing and why:
def process_image(img_name): size = (1200,1200) try: img = Image.open(img_name) img = img.filter(ImageFilter.GaussianBlur(15)) img.thumbnail(size) img.save(f'processed/{img_name}') print(f"{img_name} was processed") except Exception as e: print(f"Failed to process {img_name}: {str(e)}")
Option 2: Iterate over the results from executor.map()
Even without explicit exception handling, iterating over the returned iterator will force any hidden exceptions to surface:
def main(): t1 = time.perf_counter() with concurrent.futures.ThreadPoolExecutor() as executor: # Convert the iterator to a list to trigger all exceptions results = list(executor.map(process_image, img_names)) t2 = time.perf_counter() print(f"finished in {t2-t1} seconds")
Option 3: Stick with ProcessPoolExecutor (recommended for this task)
Image processing is CPU-intensive, and CPython's Global Interpreter Lock (GIL) limits how much parallelism you can get with threads for CPU-bound work. ProcessPoolExecutor bypasses the GIL by using separate processes, so it's actually the better choice for this kind of task anyway—it'll utilize your multi-core CPU more effectively than threads ever could.
内容的提问来源于stack exchange,提问作者DarkLeader

