Spring Data MongoDB:如何向现有GridFS文件追加二进制数据?
Great question! I’ve worked through similar TUS + GridFS implementations with Spring Data MongoDB, so let’s break down how to implement chunked appends using GridFSTemplate and Spring Data’s abstractions (since GridFS doesn’t natively support appends out of the box, but we can work with its underlying structure).
Core Background
GridFS stores files in two collections:
fs.files: Holds file metadata (filename, length, chunk size, etc.)fs.chunks: Stores binary data in fixed-size chunks (each is a separate document)
To append a TUS chunk, we need to either update an existing partial chunk or insert a new one, then update the parent file’s length in fs.files.
Step 1: Initialize the GridFS File Metadata
First, when a TUS upload starts, create the initial fs.files document (with zero length) to track the upload. You can use GridFSTemplate to access the underlying GridFSBucket for this:
import org.springframework.data.mongodb.gridfs.GridFSFile; import org.springframework.data.mongodb.gridfs.GridFSTemplate; import com.mongodb.client.gridfs.GridFSBucket; import com.mongodb.client.gridfs.model.GridFSUploadOptions; import java.io.ByteArrayInputStream; // Inject GridFSTemplate via Spring private final GridFSTemplate gridFsTemplate; public String initializeTusUpload(String filename, String contentType, long totalFileLength) { GridFSBucket gridFSBucket = gridFsTemplate.getMongoDatabase().getGridFSBucket(); // Create initial file metadata with TUS-specific fields GridFSUploadOptions uploadOptions = new GridFSUploadOptions() .metadata(new org.bson.Document() .append("tusTotalLength", totalFileLength) .append("tusOffset", 0)); // Upload an empty stream to create the fs.files entry GridFSFile initialFile = gridFSBucket.uploadFromStream( filename, new ByteArrayInputStream(new byte[0]), uploadOptions ); return initialFile.getId().toString(); // Return file ID for subsequent TUS requests }
Step 2: Append Chunks to the Existing File
For each TUS PATCH request (with offset and chunk data), calculate the target chunk index, then insert/update the chunk and update the file’s length:
import com.mongodb.client.MongoCollection; import org.bson.Document; import org.bson.types.Binary; import org.springframework.data.mongodb.core.query.Criteria; import org.springframework.data.mongodb.core.query.Query; import com.mongodb.client.model.Updates; public void appendTusChunk(String fileId, byte[] chunkData, long offset) { GridFSBucket gridFSBucket = gridFsTemplate.getMongoDatabase().getGridFSBucket(); MongoCollection<Document> chunksCollection = gridFsTemplate.getMongoDatabase() .getCollection(gridFSBucket.getChunksCollectionName()); MongoCollection<Document> filesCollection = gridFsTemplate.getMongoDatabase() .getCollection(gridFSBucket.getFilesCollectionName()); // Fetch existing file metadata GridFSFile file = gridFsTemplate.findOne(Query.query(Criteria.where("_id").is(fileId))); if (file == null) { throw new IllegalArgumentException("Upload not found: " + fileId); } int chunkSize = (int) file.getChunkSize(); int chunkIndex = (int) (offset / chunkSize); // Check if chunk already exists (for resumable uploads) Document existingChunk = chunksCollection.findOne( new Document("files_id", file.getId()).append("n", chunkIndex) ); if (existingChunk != null) { // Optional: Verify chunk integrity (e.g., checksum) before overwriting chunksCollection.updateOne( new Document("files_id", file.getId()).append("n", chunkIndex), Updates.set("data", new Binary(chunkData)) ); } else { // Insert new chunk document Document newChunk = new Document() .append("files_id", file.getId()) .append("n", chunkIndex) .append("data", new Binary(chunkData)); chunksCollection.insertOne(newChunk); } // Update file length and TUS offset long newLength = file.getLength() + chunkData.length; filesCollection.updateOne( new Document("_id", file.getId()), Updates.combine( Updates.set("length", newLength), Updates.set("metadata.tusOffset", newLength) ) ); }
Key Considerations for TUS Compliance
- Resumability: Always check for existing chunks before writing (critical for supporting interrupted uploads).
- Chunk Size: Ensure all chunks (except the final one) match the GridFS chunk size (default is 256KB, configurable via
GridFSTemplate). - Atomicity: For production, wrap chunk insertion and file metadata updates in a MongoDB transaction (requires MongoDB 4.0+) to avoid partial uploads.
- Metadata: Store TUS-specific fields (like
totalLength,offset) in the GridFS file’smetadatafield for easy access.
Why No Direct GridFSTemplate.append() Method?
Unfortunately, GridFSTemplate doesn’t expose a built-in append method—this is because GridFS was originally designed for storing complete files, not incremental appends. However, the approach above uses Spring Data’s access to the native MongoDB driver, which is fully supported and aligns with GridFS’s structure.
Comparing to Your Temporary Solution
If your temp solution involves writing chunks to disk then uploading the full file to GridFS, this approach is far more scalable: it avoids storing the entire file on server disk/memory and leverages GridFS’s native chunking for efficient storage.
内容的提问来源于stack exchange,提问作者Gustavo de Geus

