如何将文件存入二维码?Zapya实现逻辑及自研方案技术问询
Awesome question! Storing files in QR codes for cross-device transfer is a clever trick, and Zapya’s approach makes total sense once you break it down. Let’s walk through both how Zapya does it and how you can build your own version with ZXing (encoding) and BARCODE API (decoding).
QR codes have a hard limit on how much data they can hold (max ~2953 bytes for plain text in the largest QR Code version). Since most files are way bigger than that, Zapya doesn’t store the entire file in one QR code—instead, it uses a chunked approach with metadata:
- File Chunking: Split the target file into small, QR-code-sized byte chunks (usually leaving some space for metadata).
- Metadata Tagging: Each chunk gets wrapped with critical info: total number of chunks, current chunk index, file name, file size, and a hash (like MD5/SHA) for integrity checks.
- Sequential QR Codes: Generate a unique QR code for each tagged chunk. The sender displays these codes one after another (auto-switching or letting the user tap through), while the receiver scans each code to collect chunks.
- Reconstruction: Once the receiver has all chunks, it uses the metadata to reorder them correctly, verify the hash matches the original file, and save the reconstructed file to storage.
Some Zapya versions might also start with a "bootstrap QR code" that only contains the file’s core metadata (name, size, total chunks) so the receiver can pre-allocate storage and know what to expect before scanning data chunks.
Here’s a step-by-step plan using ZXing for encoding and BARCODE API for decoding:
Encoding Side (Sender Device, Using ZXing)
- Prepare the File:
- Read the target file into a byte array.
- Calculate core metadata: file name, file size, total chunks (divide file size by your chosen chunk size—aim for ~2500 bytes per chunk to leave room for metadata), and a hash of the full file (for post-transfer validation).
- Chunk and Package Data:
- Split the byte array into chunks of your chosen size.
- For each chunk, create a structured payload (JSON is easy for readability):
{ "index": 1, "total": 15, "fileName": "vacation-photos.zip", "fileSize": 37500, "fileHash": "a1b2c3d4...", "data": "base64-encoded-byte-chunk" } - Encode each chunk’s raw bytes to a Base64 string (QR codes store text, not raw binary—Base64 is the most compatible way to convert bytes to text).
- Generate QR Codes:
- Use ZXing’s
MultiFormatWriterclass to generate a QR Code for each payload string. - Display the QR codes sequentially (e.g., auto-switch every 2 seconds or let the user tap to advance) so the receiver can scan them in order.
- Use ZXing’s
Decoding Side (Receiver Device, Using BARCODE API)
- Scan and Collect Chunks:
- Use BARCODE API’s scanning functionality to read each QR code’s payload string.
- Parse the payload to extract metadata and the Base64-encoded chunk data.
- Decode the Base64 string back to raw bytes, and store each chunk with its index (use a temporary directory to hold chunks during transfer).
- Track which chunks you’ve received to avoid duplicates and know when you have all of them.
- Reconstruct and Validate:
- Once all chunks are collected, sort them by their
indexvalue. - Concatenate the sorted byte chunks to form the full file byte array.
- Calculate the hash of the reconstructed file and compare it to the
fileHashfrom the payload—if they match, the file is intact. - Write the byte array to a file with the specified
fileNamein the receiver’s storage.
- Once all chunks are collected, sort them by their
- QR Code Capacity: Stick to chunk sizes that leave enough space for metadata—don’t max out the QR code’s capacity, as this can cause encoding/decoding failures.
- Error Correction: Set a high error correction level (like QR Code’s H level, which tolerates 30% damage) to make scanning more reliable, especially if the sender’s screen is partially obscured.
- User Experience: Add progress indicators (e.g., "Scanned 5/15 chunks") on both sender and receiver to keep users informed. For large files, add a way to resume scanning if the process is interrupted (e.g., save the list of received chunks locally).
- Efficiency: Base64 adds ~33% overhead to your data. If you need more efficiency, consider using a more compact text encoding like Base85, but note that it’s less widely supported.
内容的提问来源于stack exchange,提问作者Beniamin Ionut Dobre

