如何从GCS存储的Blob文件(doc/docx/pdf)中删除元数据?
处理GCS文件上传/下载时的敏感元数据清理
核心思路
不管是上传还是下载流程,都通过流处理直接操作GCS Blob的InputStream,无需落地本地文件。核心步骤:
- 读取文件流解析为对应格式的文档对象
- 清除/覆盖敏感元数据(如作者、创建者、修改记录等)
- 将处理后的文档重新写入输出流,用于上传GCS或返回请求方
一、上传时的元数据清理流程
- 从请求中获取上传文件的InputStream(http4k可通过
request.body.stream获取) - 根据文件后缀判断类型,调用对应方法清理元数据
- 将处理后的输出流上传至GCS,替换原Blob
代码示例
1. DOCX文件处理(Apache POI)
import org.apache.poi.openxml4j.opc.OPCPackage import org.apache.poi.xwpf.usermodel.XWPFDocument import java.io.ByteArrayOutputStream fun cleanDocxMetadata(inputStream: InputStream): ByteArrayOutputStream { val outputStream = ByteArrayOutputStream() OPCPackage.open(inputStream).use { opc -> val doc = XWPFDocument(opc) val properties = opc.packageProperties // 清除敏感核心元数据 properties.creatorProperty = null properties.lastModifiedByProperty = null properties.createdProperty = null properties.modifiedProperty = null properties.titleProperty = null properties.subjectProperty = null doc.write(outputStream) } return outputStream }
2. DOC文件处理(Apache POI HWPF)
import org.apache.poi.hwpf.HWPFDocument import org.apache.poi.hwpf.usermodel.DocumentProperties import java.io.ByteArrayOutputStream fun cleanDocMetadata(inputStream: InputStream): ByteArrayOutputStream { val outputStream = ByteArrayOutputStream() HWPFDocument(inputStream).use { doc -> val properties: DocumentProperties = doc.summaryInformation // 清除敏感元数据 properties.author = "" properties.lastAuthor = "" properties.createDateTime = null properties.lastSaveDateTime = null properties.title = "" properties.subject = "" doc.write(outputStream) } return outputStream }
3. PDF文件处理(PDFBox)
import org.apache.pdfbox.pdmodel.PDDocument import org.apache.pdfbox.pdmodel.PDDocumentInformation import java.io.ByteArrayOutputStream fun cleanPdfMetadata(inputStream: InputStream): ByteArrayOutputStream { val outputStream = ByteArrayOutputStream() PDDocument.load(inputStream).use { doc -> val info: PDDocumentInformation = doc.documentInformation // 清除敏感元数据 info.author = "" info.creator = "" info.creationDate = null info.modificationDate = null info.title = "" info.subject = "" // 移除所有自定义元数据 info.customMetadataKeys.forEach { key -> info.removeCustomMetadataValue(key) } doc.save(outputStream) } return outputStream }
4. 上传到GCS的整合逻辑
// 假设已初始化GCS存储客户端 val storage = StorageOptions.getDefaultInstance().service fun uploadCleanedFile(bucketName: String, blobName: String, inputStream: InputStream, contentType: String) { val cleanedStream = when (contentType) { "application/vnd.openxmlformats-officedocument.wordprocessingml.document" -> cleanDocxMetadata(inputStream) "application/msword" -> cleanDocMetadata(inputStream) "application/pdf" -> cleanPdfMetadata(inputStream) else -> throw IllegalArgumentException("Unsupported file type") } storage.create( BlobInfo.newBuilder(bucketName, blobName).setContentType(contentType).build(), cleanedStream.toByteArray().inputStream() ) }
二、下载时的元数据清理流程
- 从GCS获取目标Blob的InputStream
- 调用上述清理方法处理元数据
- 将处理后的流直接返回给请求方(http4k可通过
Response(OK).body(cleanedStream)返回)
代码示例
fun getCleanedFileFromGcs(bucketName: String, blobName: String): ByteArrayInputStream { val blob = storage.get(bucketName, blobName) val inputStream = blob.content val contentType = blob.contentType val cleanedStream = when (contentType) { "application/vnd.openxmlformats-officedocument.wordprocessingml.document" -> cleanDocxMetadata(inputStream) "application/msword" -> cleanDocMetadata(inputStream) "application/pdf" -> cleanPdfMetadata(inputStream) else -> throw IllegalArgumentException("Unsupported file type") } return cleanedStream.toByteArray().inputStream() } // http4k路由示例 val app = routes( GET("/download/{blobName}") { req -> val blobName = req.path("blobName") val cleanedStream = getCleanedFileFromGcs("your-bucket", blobName) Response(OK).body(cleanedStream).header("Content-Disposition", "attachment; filename=$blobName") } )
关键注意事项
- 使用
use块自动关闭流资源,避免内存泄漏 - 处理大文件时,可改用
PipedInputStream/PipedOutputStream实现流式处理,避免一次性加载整个文件到内存 - GCS本身的Blob元数据(如存储的自定义元数据)需单独通过
BlobInfo修改,上述代码仅清理文件内容内部的元数据
内容的提问来源于stack exchange,提问作者landge shubham
相关产品推荐
相关产品推荐

