AWS Lambda环境读取文本文件遇文件未找到错误的解决方法
AWS Lambda 读取文件时出现 FileNotFoundError 的修复方案
问题描述
在AWS Lambda环境中尝试读取文本文件data.txt,将其转换为CSV后再转为JSON格式,但始终触发FileNotFoundError,即使使用了复制的完整路径也无法解决。
原始代码
import pandas as pd import csv import json dataframe1 = pd.read_csv(r'/data.txt', sep="|") # storing this dataframe in a csv file dataframe1.to_csv('CSV_CONVERTED.csv', index = None) def csv_to_json(event=None, context=None ): jsonArray = [] csvFilePath = r'/CSV_CONVERTED.csv' jsonFilePath = r'/data.json' #read the csv file with open(csvFilePath, encoding='utf-8') as csvf: #load csv file data using csv library's dictionary reader csvReader = csv.DictReader(csvf) #convert each csv row into python dict for row in csvReader: #add this python dict to json array jsonArray.append(row) #convert python jsonArray to JSON String and write to file with open(jsonFilePath, 'w', encoding='utf-8') as jsonf: jsonString = json.dumps(jsonArray, indent=4) jsonf.write(jsonString) return{ 'statusCode': 200, 'body': 'Success' } print(csv_to_json())
错误信息
{ "errorMessage": "[Errno 2] No such file or directory: '/data.txt'", "errorType": "FileNotFoundError", "stackTrace": [ " File \"/var/lang/lib/python3.8/imp.py\", line 234, in load_module\n return load_source(name, filename, file)\n", " File \"/var/lang/lib/python3.8/imp.py\", line 171, in load_source\n module = _load(spec)\n", " File \"\", line 702, in _load\n", " File \"\", line 671, in _load_unlocked\n", " File \"\", line 843, in exec_module\n", " File \"\", line 219, in _call_with_frames_removed\n", " File \"/var/task/Convert.py\", line 6, in \n dataframe1 = pd.read_csv(r'/data.txt', sep=\"|\")\n", " File \"/opt/python/pandas/util/_decorators.py\", line 211, in wrapper\n return func(*args, **kwargs)\n", " File \"/opt/python/pandas/util/_decorators.py\", line 331, in wrapper\n return func(*args, **kwargs)\n", " File \"/opt/python/pandas/io/parsers/readers.py\", line 950, in read_csv\n return _read(filepath_or_buffer, kwds)\n", " File \"/opt/python/pandas/io/parsers/readers.py\", line 605, in _read\n parser = TextFileReader(filepath_or_buffer, **kwds)\n", " File \"/opt/python/pandas/io/parsers/readers.py\", line 1442, in __init__\n self._engine = self._make_engine(f, self.engine)\n", " File \"/opt/python/pandas/io/parsers/readers.py\", line 1735, in _make_engine\n self.handles = get_handle(\n", " File \"/opt/python/pandas/io/common.py\", line 856, in get_handle\n handle = open(\n" ] }
问题原因
AWS Lambda的执行环境有严格的文件系统限制:
- 根目录只读:Lambda的根目录
/是只读权限,无法直接读取或写入文件。 - 部署包文件路径:如果
data.txt是和代码一起打包上传的,它会被放在Lambda的工作目录/var/task/下,而非根目录。 - 可写目录限制:Lambda中只有
/tmp目录是临时可写的,其他目录不允许写入操作。
修复方案
方案1:处理部署包内的文件
如果data.txt是和代码一起打包上传的,修改路径为工作目录路径,并将生成的文件写入/tmp目录:
import pandas as pd import csv import json # 读取部署包内的data.txt(使用相对路径或绝对路径/var/task/data.txt) dataframe1 = pd.read_csv('./data.txt', sep="|") # 将生成的CSV写入/tmp目录(唯一可写路径) temp_csv = '/tmp/CSV_CONVERTED.csv' dataframe1.to_csv(temp_csv, index=None) def csv_to_json(event=None, context=None ): jsonArray = [] csvFilePath = temp_csv # JSON文件同样写入/tmp目录 jsonFilePath = '/tmp/data.json' with open(csvFilePath, encoding='utf-8') as csvf: csvReader = csv.DictReader(csvf) for row in csvReader: jsonArray.append(row) with open(jsonFilePath, 'w', encoding='utf-8') as jsonf: jsonString = json.dumps(jsonArray, indent=4) jsonf.write(jsonString) return{ 'statusCode': 200, 'body': 'Success' } print(csv_to_json())
方案2:从S3读取文件(动态文件场景)
如果data.txt存储在S3中,需要先将文件下载到/tmp目录再处理:
import pandas as pd import csv import json import boto3 s3 = boto3.client('s3') def csv_to_json(event=None, context=None ): # 从S3下载文件到/tmp目录 s3.download_file('your-bucket-name', 'data.txt', '/tmp/data.txt') # 读取下载的文件 dataframe1 = pd.read_csv('/tmp/data.txt', sep="|") temp_csv = '/tmp/CSV_CONVERTED.csv' dataframe1.to_csv(temp_csv, index=None) jsonArray = [] csvFilePath = temp_csv jsonFilePath = '/tmp/data.json' with open(csvFilePath, encoding='utf-8') as csvf: csvReader = csv.DictReader(csvf) for row in csvReader: jsonArray.append(row) with open(jsonFilePath, 'w', encoding='utf-8') as jsonf: jsonString = json.dumps(jsonArray, indent=4) jsonf.write(jsonString) # 可选:将生成的JSON上传回S3 s3.upload_file('/tmp/data.json', 'your-bucket-name', 'data.json') return{ 'statusCode': 200, 'body': 'Success' } print(csv_to_json())
注意事项
- 打包部署时,确保
data.txt和代码文件在同一目录下,一起压缩为ZIP包上传。 /tmp目录的存储空间最大为10GB(取决于Lambda运行时版本),文件会在函数执行结束后自动清除,如需持久化需上传到S3等存储服务。- 如果不确定文件路径,可以在代码中加入
import os; print(os.getcwd())打印当前工作目录,确认部署包文件的位置。
内容的提问来源于stack exchange,提问作者Afridi Rahman
相关产品推荐
相关产品推荐

