如何在Python中对生成器式文本流进行压缩与加密?
Great question! The good news is pure streaming compression absolutely exists—you don’t need to load your entire generator output into memory to compress it. Let’s break down the solutions, including native streaming approaches and workarounds if you run into edge cases:
Most common compression algorithms are designed to support incremental, streaming processing. The key ones you can use out of the box include:
- Deflate/zlib/gzip: The standard for web and general-purpose compression; uses a sliding window to process data in chunks without needing the full dataset.
- Brotli: Google’s modern compression algorithm, optimized for web content, with full streaming support.
- LZ4: A fast, low-overhead algorithm with dedicated streaming APIs for both compression and decompression.
Example: Python Streaming Compression + Encryption
Here’s a practical implementation using Python’s built-in zlib (for streaming compression) and cryptography (for streaming AES-GCM encryption):
import zlib from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes from cryptography.hazmat.backends import default_backend def stream_process(generator): # Initialize streaming compressor (zlib) compressor = zlib.compressobj(level=zlib.Z_BEST_COMPRESSION) # Initialize streaming AES-GCM encryptor (choose a secure key/IV in production!) key = b"your-32-byte-secret-key-here" # AES-256 requires 32-byte key iv = b"your-12-byte-iv-here" # GCM recommends 12-byte IV cipher = Cipher(algorithms.AES(key), modes.GCM(iv), backend=default_backend()) encryptor = cipher.encryptor() # Process each chunk from the generator for chunk in generator: # Compress the current chunk (outputs data immediately if possible) compressed_chunk = compressor.compress(chunk) if compressed_chunk: # Encrypt the compressed chunk and yield it yield encryptor.update(compressed_chunk) # Flush any remaining data from the compressor final_compressed = compressor.flush() if final_compressed: yield encryptor.update(final_compressed) # Finalize encryption and yield the GCM authentication tag yield encryptor.finalize() + encryptor.tag
This code processes each chunk from your generator as it arrives—no full dataset is ever stored in memory. The compressor maintains a small sliding window of historical data, and the encryptor processes bytes incrementally.
Algorithms like Deflate and Brotli rely on sliding window compression: they only need to keep track of a recent segment of data (the "window") to find repeated patterns, not the entire dataset. This makes them inherently compatible with streaming workflows.
If you’re working with a compression algorithm that lacks native streaming support (rare, but possible with some niche high-compression tools), you can use chunked batch processing as a near-streaming alternative:
- Split your generator’s output into fixed-size chunks (e.g., 64KB or 1MB).
- Compress each chunk individually (or accumulate chunks until a threshold is met, then compress).
- Encrypt each compressed chunk and add a small header to mark chunk boundaries (so the decompressor knows where one chunk ends and the next begins).
This approach keeps memory usage low (only one chunk is loaded at a time) and mimics the behavior of a true stream.
- Always compress first, then encrypt: Encrypted data is pseudorandom, so compressing it after encryption will do nothing (or even increase file size). Compression works best on structured, non-random data.
- Choose streaming-friendly encryption modes: Use modes like GCM or CTR (not ECB, which is insecure, or CBC, which requires block alignment and is less straightforward for streaming).
- Save state if needed: If you need to pause/resume processing, you’ll need to serialize the compressor and encryptor’s internal state (e.g., zlib’s compression state, GCM’s counter) to pick up where you left off.
内容的提问来源于stack exchange,提问作者user1827975

