You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用DuckDB读取S3中Parquet文件时遇段错误核心转储问题排查

解决DuckDB读取S3上Parquet文件时的Segmentation Fault问题

问题背景

每日有数千个同结构的Parquet文件存入S3存储桶,我用Python3结合DuckDB扩展读取这些文件并提取数据子集。

使用的代码片段如下:

dtt = [datetime数组]

for dt in dtt:
    print('processing for date :' + str(dt))
    ddt=dt.date()
    dstr=str(dt)
    yr=dstr[0:4]
    mn=dstr[5:7]
    dy=dstr[8:10]

    rurl = "s3://s3bucket/hind/"+str(yr) +'/'+str(mn)+'/'+str(dy)+'/*.parquet'
    nfdf=duckdb.sql("select maid,date,hour,h3id from read_parquet('" + str(rurl) +"') where h3id In " + str(tuple(uhids))).df()

每个Parquet文件大小为25-35MB,上述代码并非每次都能正常运行,会出现Segmentation Fault Core Dumped错误,且无规律可循,无法定位原因。

我尝试过多种AWS实例类型:m6a.4xlarge最稳定但速度极慢;使用m6a.24xlarge或m6a.48xlarge时速度更快,但会很快触发核心转储错误。实例无内存或磁盘空间不足问题,也无并行进程运行,处理器未处于繁忙状态。

当前环境:AWS Cli 2.13.25、Python 3.8.10、Ubuntu 22.04。求解决方法,更换方案也可接受。

参考修改方案(按Dean建议调整)

将原代码中nfdf=duckdb.sql("select....部分替换为:

tempfil=str('p1_')+str(p)
con=duckdb.connect('file.db')
con.sql('drop table ' + str(tempfil))
con.sql('CREATE TABLE '+str(tempfil)+'(maid text,date date,hour integer,h3id varchar)')
con.sql("insert into " + str(tempfil) + " select maid,date,hour,h3id from read_parquet('" + str(rurl) +"') where h3id In " + str(tuple(uhids)))
ifile="/home/ubuntu/tdata1/p3_"+str(p)+"_"+str(dy)+str(mn)+str(yr)+".parquet"
con.sql("COPY "+ str(tempfil)  +" TO '" + str(ifile)  +"'")
con.table(tempfil).show()
con.sql('drop table ' + str(tempfil))
con.close()

配置测试报错情况

按@Carlo建议使用以下代码测试:

import duckdb
duckdb.execute("SET home_directory='/home/ubuntu/'")
duckdb.execute("INSTALL httpfs")
duckdb.execute("LOAD httpfs")
duckdb.execute("SET s3_access_key_id='XXXXXXXXX'")
duckdb.execute("SET s3_secret_access_key='XXXXXXXXXXX'")
duckdb.execute("SET s3_region='XXXXXXXX'")

duckdb.sql('PRAGMA platform')
duckdb.sql('PRAGMA version')
duckdb.sql("FROM duckdb_extensions() WHERE extension_name == 'httpfs'")

运行后出现报错:

Traceback (most recent call last):
  File "ddbconfig.py", line 10, in <module>
    duckdb.sql('PRAGMA platform')
duckdb.CatalogException: Catalog Error: Pragma Function with name platform does not exist!
Did you mean "tpch"?

内容的提问来源于stack exchange,提问作者Apricot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 17:33:15