如何在上传文件至Azure Blob Storage容器时添加元数据/标签,并在Blob Trigger中利用其进行文件筛选(含路径场景示例)
Perfect question—let's walk through exactly how to add metadata/tags when uploading blobs to Azure Blob Storage, then how to filter for specific ones (like your DocType=RequiresScan scenario) in a Blob Trigger, including path-based filters too.
First, it's important to note that metadata and tags are distinct in Azure Blob Storage:
- Metadata: Unindexed key-value pairs attached to the blob, ideal for storing blob-specific properties. Keys are automatically stored in lowercase.
- Tags: Indexed key-value pairs, optimized for querying and filtering blobs at scale.
Uploading with Metadata
Here's how to add metadata during upload using the .NET SDK (similar patterns work for Python, Azure CLI, and other tools):
using Azure.Storage.Blobs; using System.IO; using System.Collections.Generic; var connectionString = "your-storage-connection-string"; var containerName = "your-container"; var blobPath = "documents/invoice.pdf"; var blobServiceClient = new BlobServiceClient(connectionString); var containerClient = blobServiceClient.GetBlobContainerClient(containerName); var blobClient = containerClient.GetBlobClient(blobPath); // Define your metadata key-value pairs var metadata = new Dictionary<string, string> { { "DocType", "RequiresScan" }, { "UploadDate", "2024-05-20" }, { "UploadedBy", "FinanceTeam" } }; // Upload the file with metadata attached using var fileStream = File.OpenRead("local/path/to/invoice.pdf"); await blobClient.UploadAsync(fileStream, new BlobUploadOptions { Metadata = metadata });
Uploading with Tags
Tags can be added either during upload or after the blob is stored. Here's the upload-time approach:
// Define your indexed tags var tags = new Dictionary<string, string> { { "DocType", "RequiresScan" }, { "Department", "Finance" }, { "Status", "Pending" } }; // Upload with tags included await blobClient.UploadAsync(fileStream, new BlobUploadOptions { Tags = tags }); // Alternatively, set tags after upload: // await blobClient.SetTagsAsync(tags);
Blob Triggers don't support direct filtering on metadata/tags via the binding itself, but you can handle filtering in two ways: post-trigger checks (simple, but triggers all blobs first) or pre-trigger filtering with Event Grid (more efficient, only triggers matching blobs).
1. Filtering by Metadata (Post-Trigger)
Inject the blob's metadata into your function and validate the DocType=RequiresScan pair. Remember metadata keys are stored in lowercase:
using Microsoft.Azure.WebJobs; using Microsoft.Extensions.Logging; using System.Collections.Generic; using System.IO; namespace BlobTriggerMetadataFilter { public static class ProcessScannedDocuments { [FunctionName("ProcessScannedDocuments")] public static async Task Run( [BlobTrigger("your-container/{name}", Connection = "AzureWebJobsStorage")] Stream myBlob, string name, IDictionary<string, string> metadata, // Injects blob metadata ILogger log) { // Check for the required metadata (use lowercase key) if (metadata.TryGetValue("doctype", out var docType) && docType.Equals("RequiresScan", System.StringComparison.OrdinalIgnoreCase)) { log.LogInformation($"Processing document requiring scan: {name}"); // Add your scan/processing logic here await ProcessDocument(myBlob); } else { log.LogInformation($"Skipping {name} - does not match DocType=RequiresScan"); return; // Exit early if no match } } private static async Task ProcessDocument(Stream blobStream) { // Your custom processing logic goes here } } }
2. Filtering by Tags (Post-Trigger or Pre-Trigger)
Post-Trigger Check
Inject a BlobClient to fetch tags and validate the criteria:
[FunctionName("ProcessTaggedDocuments")] public static async Task Run( [BlobTrigger("your-container/{name}", Connection = "AzureWebJobsStorage")] Stream myBlob, string name, BlobClient blobClient, // Injects BlobClient to fetch tags ILogger log) { var tagResponse = await blobClient.GetTagsAsync(); var tags = tagResponse.Value.Tags; if (tags.TryGetValue("DocType", out var docType) && docType.Equals("RequiresScan", System.StringComparison.OrdinalIgnoreCase)) { log.LogInformation($"Processing tagged document: {name}"); await ProcessDocument(myBlob); } else { log.LogInformation($"Skipping {name} - incorrect tags"); return; } }
Pre-Trigger Filtering with Event Grid (More Efficient)
To avoid triggering the function for non-matching blobs, swap the Blob Trigger for an Event Grid Trigger. When setting up your Event Grid subscription:
- Set the topic to your storage account's blob events
- Add these filters:
- Subject filter:
Begins with /blobServices/default/containers/your-container/blobs/ - Advanced filter:
data.tags.DocType String equals RequiresScan
- Subject filter:
This way, only blobs with the correct tag will send an event to your function, reducing unnecessary executions.
3. Path-Based Filtering
If you need to filter by folder structure or file type, define this directly in the Blob Trigger path:
- Trigger only blobs in a specific folder:
[BlobTrigger("your-container/scanned-documents/{name}", Connection = "AzureWebJobsStorage")] - Trigger only PDF files in any subfolder:
[BlobTrigger("your-container/**/*.pdf", Connection = "AzureWebJobsStorage")]
Combine Path + Metadata/Tag Filtering
You can easily combine path filters with metadata/tag checks for granular control:
[FunctionName("ProcessFilteredDocuments")] public static async Task Run( [BlobTrigger("your-container/pending-scans/{name}", Connection = "AzureWebJobsStorage")] Stream myBlob, string name, IDictionary<string, string> metadata, ILogger log) { if (metadata.TryGetValue("doctype", out var docType) && docType.Equals("RequiresScan", System.StringComparison.OrdinalIgnoreCase)) { log.LogInformation($"Processing {name} from pending-scans folder with correct metadata"); await ProcessDocument(myBlob); } else { log.LogInformation($"Skipping {name} - does not meet criteria"); return; } }
- Use metadata for blob-specific properties that don't need frequent querying.
- Use tags when you need efficient, scalable filtering (pair with Event Grid for pre-trigger filtering to save resources).
- Path-based filtering is the simplest way to narrow down blobs by folder or file type.
内容的提问来源于stack exchange,提问作者user15878634

