寻求免重复压缩的工具/库:从Zip生成含gzip版本的内容文件夹
Great question—you’ve nailed a key optimization here: reusing the existing deflate compression streams from ZIP files eliminates the redundant decompress-recompress cycle, which will save a ton of CPU time for large files. Let’s cover your options across Linux command-line tools and the languages you mentioned:
Linux Command-Line Approach
There’s no single out-of-the-box tool that does this directly, but you can build a script using existing utilities to parse ZIP structures and repackage deflate streams into valid GZIP files. Here’s a breakdown:
Parse ZIP Structure with
zipdetails
First, installzipdetails(a Perl-based tool for inspecting ZIP internals) to locate the deflate data blocks for each file. Run:zipdetails large_file.zip | grep -A5 -B5 "DEFLATED"This will show you the offset and length of each deflated stream, along with metadata like CRC32 and uncompressed file size.
Extract and Repackage with
dd+ Shell Scripting
For each deflated entry, useddto extract the raw deflate data, then prepend the GZIP header and append the CRC32/uncompressed size footer. A simplified snippet for one file might look like:# GZIP header (ID: 1f8b, Compression method: 8 (deflate), flags: 0, mtime: 0, OS: 3 (Unix)) printf '\x1f\x8b\x08\x00\x00\x00\x00\x00\x00\x03' > output/file.txt.gz # Extract deflate stream from ZIP (adjust offset and count to match zipdetails output) dd if=large_file.zip skip=$OFFSET bs=1 count=$DEFLATE_LENGTH >> output/file.txt.gz # Append CRC32 and uncompressed size (little-endian) printf '\x%s\x%s\x%s\x%s' $(printf '%08x' $CRC32 | sed 's/../\\x&/g') >> output/file.txt.gz printf '\x%s\x%s\x%s\x%s' $(printf '%08x' $UNCOMPRESSED_SIZE | sed 's/../\\x&/g') >> output/file.txt.gz # Extract the original uncompressed file for your output folder unzip -j large_file.zip file.txt -d output/You’d wrap this in a loop to process all entries in the ZIP.
Alternative: Use
bsdtar(Libarchive) for Simplified Extraction
Whilebsdtardoesn’t reuse deflate streams directly, it’s faster than standardunzip+gzipfor large files. If you need a quick solution without writing a full script, it’s a good middle ground:# Extract uncompressed files to output/ bsdtar -xf large_file.zip -C output/ # Compress with pigz (parallel gzip) for speed (still recompresses, but faster) find output/ -type f -exec pigz -k {} \;Note: This still recompresses, but it’s a faster fallback if the stream-reuse script is too complex.
Language-Specific Libraries
Ruby
Use the rubyzip gem to access ZIP entries directly and repackage deflate streams into GZIP files. Here’s a code example:
require 'zip' require 'zlib' input_zip = 'large_file.zip' output_dir = 'output' Dir.mkdir(output_dir) unless Dir.exist?(output_dir) Zip::File.open(input_zip) do |zip_file| zip_file.each do |entry| next if entry.directory? # Extract original uncompressed file entry.extract(File.join(output_dir, entry.name)) # Skip if entry isn't using deflate compression next unless entry.compression_method == Zip::Entry::DEFLATED # Build GZIP file gzip_path = File.join(output_dir, "#{entry.name}.gz") File.open(gzip_path, 'wb') do |f| # GZIP header f.write([0x1f8b, 8, 0, 0, 0, 0, 0, 3].pack('SCCLCCC')) # Write raw deflate stream f.write(entry.get_raw_compressed_data) # Write CRC32 and uncompressed size (little-endian) f.write([entry.crc32, entry.size].pack('LL')) end end end
Install the gem first with gem install rubyzip.
Node.js
Use yauzl (a low-level ZIP parser) to access deflate streams, then construct valid GZIP files manually. Example:
const yauzl = require('yauzl'); const fs = require('fs'); const path = require('path'); const inputZip = 'large_file.zip'; const outputDir = 'output'; fs.mkdirSync(outputDir, { recursive: true }); yauzl.open(inputZip, { lazyEntries: true }, (err, zipfile) => { if (err) throw err; zipfile.readEntry(); zipfile.on('entry', (entry) => { if (/\/$/.test(entry.fileName)) { zipfile.readEntry(); return; } // Extract original uncompressed file zipfile.openReadStream(entry, (err, readStream) => { if (err) throw err; const outputPath = path.join(outputDir, entry.fileName); readStream.pipe(fs.createWriteStream(outputPath)).on('finish', () => { // Process deflate stream if applicable if (entry.compressionMethod === 8) { // 8 = DEFLATE const gzipPath = `${outputPath}.gz`; const gzipStream = fs.createWriteStream(gzipPath); // Write GZIP header gzipStream.write(Buffer.from([ 0x1f, 0x8b, // ID 0x08, // Compression method (deflate) 0x00, // Flags 0x00, 0x00, 0x00, 0x00, // MTIME (0) 0x00, // XFL 0x03 // OS (Unix) ])); // Write raw deflate data (reopen stream to get compressed data) zipfile.openReadStream(entry, { decompress: false }, (err, compStream) => { if (err) throw err; compStream.pipe(gzipStream, { end: false }); compStream.on('end', () => { // Write CRC32 and uncompressed size (little-endian) const crcBuf = Buffer.alloc(4); crcBuf.writeUInt32LE(entry.crc32, 0); const sizeBuf = Buffer.alloc(4); sizeBuf.writeUInt32LE(entry.uncompressedSize, 0); gzipStream.write(crcBuf); gzipStream.write(sizeBuf); gzipStream.end(); }); }); } zipfile.readEntry(); }); }); }); });
Install dependencies with npm install yauzl.
C++
Use libzip to parse ZIP files and repackage deflate streams. Here’s a minimal example:
#include <zip.h> #include <fstream> #include <cstdint> #include <sys/stat.h> int main() { const char* input_zip = "large_file.zip"; const char* output_dir = "output"; mkdir(output_dir, 0755); zip_t* zf = zip_open(input_zip, ZIP_RDONLY, nullptr); if (!zf) return 1; int num_entries = zip_get_num_entries(zf, 0); for (int i = 0; i < num_entries; ++i) { const char* name = zip_get_name(zf, i, 0); if (!name) continue; // Extract uncompressed file zip_file_t* zfile = zip_fopen_index(zf, i, 0); if (!zfile) continue; std::ofstream orig_file(std::string(output_dir) + "/" + name, std::ios::binary); char buf[4096]; int bytes_read; while ((bytes_read = zip_fread(zfile, buf, sizeof(buf))) > 0) { orig_file.write(buf, bytes_read); } zip_fclose(zfile); orig_file.close(); // Check if entry uses deflate zip_stat_t stat; zip_stat_index(zf, i, 0, &stat); if (stat.comp_method != ZIP_CM_DEFLATE) continue; // Create GZIP file std::ofstream gzip_file(std::string(output_dir) + "/" + name + ".gz", std::ios::binary); // Write GZIP header uint8_t header[] = {0x1f, 0x8b, 0x08, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x03}; gzip_file.write(reinterpret_cast<char*>(header), sizeof(header)); // Write raw deflate stream zfile = zip_fopen_index(zf, i, ZIP_FL_COMPRESSED); if (!zfile) continue; while ((bytes_read = zip_fread(zfile, buf, sizeof(buf))) > 0) { gzip_file.write(buf, bytes_read); } zip_fclose(zfile); // Write CRC32 and uncompressed size (little-endian) uint32_t crc = stat.crc; uint32_t size = stat.size; gzip_file.write(reinterpret_cast<char*>(&crc), sizeof(crc)); gzip_file.write(reinterpret_cast<char*>(&size), sizeof(size)); gzip_file.close(); } zip_close(zf); return 0; }
Compile with g++ -o zip2gzip zip2gzip.cpp -lzip (ensure libzip-dev is installed).
内容的提问来源于stack exchange,提问作者mixture

