You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Lambda触发S3提取文件ID匹配并复制对应元数据文件

Solution using AWS Lambda and S3

1. Set Up S3 Trigger for Lambda

First, configure your source S3 bucket (where .mxf files are uploaded) to trigger the Lambda function whenever a new .mxf file is added:

  • Go to your source bucket → Properties → Event notifications.
  • Create a new notification:
    • Name it something like MXF_Upload_Trigger
    • Check "All object create events" under Event types
    • Add .mxf as the suffix to only trigger on .mxf file uploads
    • Select "Lambda function" as the destination, then pick your Lambda function (create the function first if it doesn’t exist)

2. Lambda Function Implementation (Python)

IAM Permissions

Your Lambda execution role needs these permissions (attach a custom policy to the role):

  • s3:GetObject for the source bucket
  • s3:ListBucket and s3:GetObject for the metadata bucket
  • s3:PutObject for the destination bucket (where you want to copy the metadata file; can be the source bucket or a separate one)

Code

import boto3
import os

s3 = boto3.client('s3')

# Pull bucket names from Lambda environment variables
METADATA_BUCKET = os.environ['METADATA_BUCKET']
# Default to source bucket if destination isn't specified
DESTINATION_BUCKET = os.environ.get('DESTINATION_BUCKET')

def lambda_handler(event, context):
    for record in event['Records']:
        source_bucket = record['s3']['bucket']['name']
        mxf_key = record['s3']['object']['key']
        
        # Skip non-.mxf files (redundant but safe)
        if not mxf_key.endswith('.mxf'):
            continue
        
        # Extract content ID from filename
        try:
            filename_segments = mxf_key.split('_')
            if len(filename_segments) < 3:
                print(f"Filename {mxf_key} doesn't follow expected pattern")
                continue
            content_id = filename_segments[2]
            print(f"Extracted content ID: {content_id}")
        except Exception as e:
            print(f"Failed to extract content ID: {str(e)}")
            continue
        
        # Find matching XML files in metadata bucket
        try:
            paginator = s3.get_paginator('list_objects_v2')
            xml_objects = paginator.paginate(
                Bucket=METADATA_BUCKET,
                Suffix='.xml'
                # Add a Prefix here if metadata files are in a specific subfolder
            )
            
            matching_xmls = []
            for page in xml_objects:
                if 'Contents' not in page:
                    continue
                for obj in page['Contents']:
                    if content_id in obj['Key']:
                        matching_xmls.append(obj['Key'])
            
            if not matching_xmls:
                print(f"No metadata files found for ID {content_id}")
                continue
            
            # Copy each matching XML to the destination bucket
            dest_bucket = DESTINATION_BUCKET or source_bucket
            for xml_key in matching_xmls:
                # Define destination path (adjust this to fit your needs)
                dest_key = f"{os.path.dirname(mxf_key)}/{os.path.basename(xml_key)}"
                
                try:
                    s3.copy_object(
                        Bucket=dest_bucket,
                        Key=dest_key,
                        CopySource={'Bucket': METADATA_BUCKET, 'Key': xml_key}
                    )
                    print(f"Copied {xml_key} to {dest_bucket}/{dest_key}")
                except Exception as e:
                    print(f"Failed to copy {xml_key}: {str(e)}")
        
        except Exception as e:
            print(f"Error accessing metadata bucket: {str(e)}")
            continue

Configure Environment Variables

In your Lambda function settings, add these environment variables:

  • METADATA_BUCKET: Name of the bucket storing your metadata XML files
  • DESTINATION_BUCKET (optional): Name of the bucket where you want to copy the metadata files. If omitted, files will be copied to the same bucket as the uploaded .mxf.

3. Test the Workflow

  • Upload an .mxf file matching your filename pattern to the source bucket
  • Check Lambda’s CloudWatch logs to confirm the content ID was extracted and the metadata file was copied
  • Verify the metadata XML appears in your destination bucket

Key Notes

  • If your metadata bucket has thousands of files, listing all XMLs every time might be slow. Consider organizing metadata files into subfolders by content ID, or using S3 Inventory to pre-index files for faster lookups.
  • The code copies all matching XML files for a content ID. If you only need the first match, modify the code to break after the first copy.
  • Double-check IAM permissions if you run into access errors—Lambda needs explicit permission for all buckets involved.

内容的提问来源于stack exchange,提问作者mpadeti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 14:47:42