如何用Python读取S3桶中匹配命名模式的文件并合并为Pandas DataFrame
读取S3中匹配模式的多个CSV并合并为DataFrame
方法一:使用Pandas内置通配符匹配(最简方案)
如果目标目录下仅存在符合fname_mmyy.csv格式的文件,直接通过通配符批量获取文件并合并:
import pandas as pd aws_credentials = { "key": "xxxx", "secret": "xxxx" } # 获取所有匹配的S3文件路径,批量读取后合并为单个DataFrame df_combined = pd.concat( pd.read_csv(file_path, storage_options=aws_credentials, encoding='latin-1') for file_path in pd.io.common.get_filenames("s3://dir/ABC/fname_*.csv", storage_options=aws_credentials) )
方法二:正则精确筛选文件(避免误读无关文件)
如果目录下包含其他fname_开头的无关文件,用正则表达式精确匹配fname_mmyy.csv格式(mm为01-12的月份,yy为两位年份):
import pandas as pd import s3fs import re aws_credentials = { "key": "xxxx", "secret": "xxxx" } # 初始化S3文件系统客户端 fs = s3fs.S3FileSystem(key=aws_credentials['key'], secret=aws_credentials['secret']) # 列出目标目录下的所有文件 all_files = fs.ls("dir/ABC") # 用正则筛选符合格式的文件路径 file_pattern = re.compile(r"fname_([0-1][0-9][0-9]{2})\.csv$") target_files = [f"s3://{file}" for file in all_files if file_pattern.search(file)] # 批量读取并合并DataFrame df_combined = pd.concat( pd.read_csv(file, storage_options=aws_credentials, encoding='latin-1') for file in target_files )
注意事项
- 确保所有目标文件的列结构完全一致,否则合并会出现列名不匹配或缺失值问题。
- 若文件体积过大,可在
pd.read_csv中添加chunksize参数分块读取,避免内存溢出。
内容的提问来源于stack exchange,提问作者kgh
相关产品推荐
相关产品推荐

