PySpark加载多分区ORC文件遇load参数超限,求替代方案
解决Spark加载多个ORC文件时
TypeError: load() takes at most 4 arguments的问题 问题背景
我尝试一次性加载多个分区ORC文件,单个文件加载正常,但加载24个文件时触发错误:TypeError: load() takes at most 4 arguments (24 given)。除了加载后执行union操作外,没找到相关限制文档及其他临时解决办法,想问问有没有更优的替代方案?
复现代码
basepath = '/file/' paths = ['/file/df201601.orc', '/file/df201602.orc', '/file/df201603.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', '/file/df201604.orc', '/file/df201605.orc', '/file/df201606.orc', ] df = sqlContext.read.format('orc') \ .options(header='true',inferschema='true',basePath=basepath)\ .load(*paths)
报错信息
TypeError Traceback (most recent call last) <ipython-input-43-7fb8fade5e19> in <module>() ---> 37 df = sqlContext.read.format('orc') .options(header='true', inferschema='true',basePath=basepath) .load(*paths) 38 TypeError: load() takes at most 4 arguments (24 given)
解决方案
别担心,这个问题我之前也碰到过,其实是Spark Python API里load()方法的参数设计小坑——它的底层实现限制了位置参数的数量,但我们有几个比手动union优雅得多的方案:
方案1:直接传入路径列表(不用解包)
Spark的load()方法其实原生支持接收一个路径列表作为参数,根本不需要用*来解包。修改代码如下:
df = sqlContext.read.format('orc') \ .options(header='true', inferschema='true', basePath=basepath)\ .load(paths)
这种方式下Spark会自动批量读取所有指定路径的ORC文件,内部会做合并优化,性能比手动union好很多。
方案2:用通配符匹配路径
如果你的文件命名有规律(比如都是df2016开头的ORC文件),直接用通配符匹配会更简洁:
df = sqlContext.read.format('orc') \ .options(header='true', inferschema='true', basePath=basepath)\ .load('/file/df2016*.orc')
方案3:直接读取整个目录
如果所有目标ORC文件都在/file/目录下,且没有其他无关文件干扰,直接加载目录路径就行:
df = sqlContext.read.format('orc') \ .options(header='true', inferschema='true', basePath=basepath)\ .load(basepath)
补充说明
- 为什么解包会报错?因为Spark Python API的
load()虽然定义了*paths参数,但底层的Java API对接时限制了位置参数数量,直接传列表才是官方推荐的批量读取方式。 - 以上方案都比手动union多个DataFrame高效,因为Spark可以在读取阶段就进行数据合并的优化,避免多次IO和合并的额外开销。
内容的提问来源于stack exchange,提问作者E B
相关产品推荐
相关产品推荐

