Spark Mongo-Connector读取DocumentDB失败,如何关闭collStats功能?
解决Spark Mongo-Connector连接DocumentDB时collStats命令报错的问题
问题根源是Spark Mongo-Connector默认会通过collStats命令获取集合统计信息来做数据分区,但DocumentDB不支持这个命令。printSchema()不需要执行分区逻辑所以能正常返回,而show()会触发分区操作,因此报错。
解决方法是强制连接器使用单分区模式,避免调用collStats命令,具体是添加spark.mongodb.read.partitioner配置项,指定为MongoSinglePartitioner:
修改后的代码示例:
dataFrame = spark.read.format("mongodb") \ .option("spark.mongodb.database", "testdb") \ .option("spark.mongodb.collection", "collection1") \ .option("spark.mongodb.read.partitioner", "MongoSinglePartitioner") \ .load()
这个配置会让连接器用单个分区读取数据,完全跳过依赖collStats的分区逻辑,就能正常执行show()操作了。
如果后续需要多分区读取,也可以尝试MongoRangePartitioner,但需要手动指定分区键(比如_id)和分区数量,这种方式不依赖collStats,也能在DocumentDB上工作:
dataFrame = spark.read.format("mongodb") \ .option("spark.mongodb.database", "testdb") \ .option("spark.mongodb.collection", "collection1") \ .option("spark.mongodb.read.partitioner", "MongoRangePartitioner") \ .option("spark.mongodb.read.partitioner.options.partitionKey", "_id") \ .option("spark.mongodb.read.partitioner.options.numberOfPartitions", "4") \ .load()
内容的提问来源于stack exchange,提问作者Jonathanv
相关产品推荐
相关产品推荐

