AWS Glue作业读取S3 JSON文件时出现FileNotFoundError问题
问题原因
Python内置的open()函数仅支持访问本地文件系统路径,无法直接识别s3://协议的远程存储路径。而Spark的spark.read.csv等API是AWS Glue/Spark专门适配过S3存储的,因此可以直接读取S3上的文件。权限配置和路径调整无法解决这个底层的API适配问题。
解决方案
以下几种方法可以解决这个问题:
方法1:用boto3直接读取S3文件内容
利用AWS SDK for Python(boto3)读取S3对象的内容,再解析JSON:
import boto3 import json # 初始化S3客户端 s3_client = boto3.client('s3') # 定义S3桶名和文件键 bucket = 'bucket-name' file_key = 'configs/mapping.json' # 获取文件内容并解析 response = s3_client.get_object(Bucket=bucket, Key=file_key) json_str = response['Body'].read().decode('utf-8') mapping_config = json.loads(json_str)
方法2:用Spark API读取JSON后转成Python对象
如果JSON文件格式符合Spark的读取要求,可以用Spark读取后转换为Python字典:
import json from pyspark.sql import SparkSession spark = SparkSession.builder.getOrCreate() # 读取S3上的JSON文件为DataFrame json_df = spark.read.json('s3://bucket-name/configs/mapping.json') # 转换为Python字典(适用于单条JSON记录的文件) mapping_config = json_df.toPandas().to_dict('records')[0]
方法3:下载S3文件到本地临时目录后用open()读取
Glue作业提供了/tmp/临时目录,可以先把S3文件下载到本地,再用open()读取:
import boto3 import json import os s3_client = boto3.client('s3') bucket = 'bucket-name' file_key = 'configs/mapping.json' local_temp_path = '/tmp/mapping.json' # 下载文件到本地临时目录 s3_client.download_file(bucket, file_key, local_temp_path) # 用open()读取本地文件 with open(local_temp_path, 'r') as f: mapping_config = json.load(f) # 可选:清理临时文件 os.remove(local_temp_path)
内容的提问来源于stack exchange,提问作者J.A.
相关产品推荐
相关产品推荐

