You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取S3桶中匹配命名模式的文件并合并为Pandas DataFrame

读取S3中匹配模式的多个CSV并合并为DataFrame

方法一:使用Pandas内置通配符匹配(最简方案)

如果目标目录下仅存在符合fname_mmyy.csv格式的文件,直接通过通配符批量获取文件并合并:

import pandas as pd

aws_credentials = { 
    "key": "xxxx", 
    "secret": "xxxx" 
}

# 获取所有匹配的S3文件路径,批量读取后合并为单个DataFrame
df_combined = pd.concat(
    pd.read_csv(file_path, storage_options=aws_credentials, encoding='latin-1')
    for file_path in pd.io.common.get_filenames("s3://dir/ABC/fname_*.csv", storage_options=aws_credentials)
)

方法二:正则精确筛选文件(避免误读无关文件)

如果目录下包含其他fname_开头的无关文件,用正则表达式精确匹配fname_mmyy.csv格式(mm为01-12的月份,yy为两位年份):

import pandas as pd
import s3fs
import re

aws_credentials = { 
    "key": "xxxx", 
    "secret": "xxxx" 
}

# 初始化S3文件系统客户端
fs = s3fs.S3FileSystem(key=aws_credentials['key'], secret=aws_credentials['secret'])

# 列出目标目录下的所有文件
all_files = fs.ls("dir/ABC")

# 用正则筛选符合格式的文件路径
file_pattern = re.compile(r"fname_([0-1][0-9][0-9]{2})\.csv$")
target_files = [f"s3://{file}" for file in all_files if file_pattern.search(file)]

# 批量读取并合并DataFrame
df_combined = pd.concat(
    pd.read_csv(file, storage_options=aws_credentials, encoding='latin-1')
    for file in target_files
)

注意事项

  • 确保所有目标文件的列结构完全一致,否则合并会出现列名不匹配或缺失值问题。
  • 若文件体积过大,可在pd.read_csv中添加chunksize参数分块读取,避免内存溢出。

内容的提问来源于stack exchange,提问作者kgh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 18:50:28