Fabric Notebook运行PySpark时遇Py4JJavaError:底层路径不存在
Microsoft Fabric PySpark Notebook执行报错排查
错误信息
Py4JJavaError An error occurred while calling o6734.collectToPython. : org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 18.0 failed 4 times, most recent failure: Lost task 0.3 in stage 18.0 (TID 23) (vm-4f039835 executor 1): com.microsoft.sqlserver.jdbc.SQLServerException: An error occurred during the current command (Done status 0). Failed to complete the command because the underlying location does not exist. Underlying data description: table '\<lakehouse table path\>', file '\<lakehouse table url\>'.
相关代码
import com.microsoft.spark.fabric from com.microsoft.spark.fabric.Constants import Constants spark = SparkSession.builder.appName("create_availability").getOrCreate() # Build SQL query to find min and max date from source view query = "SELECT MIN(week_start_date) AS earliest_date, MAX(week_start_date) AS latest_date FROM schema.source_view" # Load from SQL endpoint table_load_spark = ( spark.read.option(Constants.WorkspaceId, WS_ID) .option(Constants.DatabaseName, LAKEHOUSE_NAME) .synapsesql(query) ) dates = table_load_spark.first()
问题背景与排查方向
- 报错触发点:调用
first()方法时,collectToPython执行失败 - 场景特殊性:仅通过服务主体调度的流水线运行时出现该错误,手动运行Notebook无异常
- 权限现状:服务主体已拥有对应Lakehouse和SQL端点的权限,且相同权限配置在其他工作区可正常运行
- 推测原因:Spark任务在执行查询时,间接访问了源视图所在架构下服务主体无权限访问的其他表,导致底层数据路径无法访问
内容的提问来源于stack exchange,提问作者Bekka C
相关产品推荐
相关产品推荐

