Scala读取含//的文件路径失败,如何正确处理特殊字符?
问题分析与解决办法
首先要明确:Scala的Source.fromFile只能读取本地文件系统(或已挂载到本地的存储)的文件,它基于JVM的File API实现,根本不支持dfss://这类分布式存储协议的路径——这才是你报错的核心原因,和路径里的//被识别为单个/没关系。
可行解决方案
方案1:使用对应分布式存储的官方SDK读取
从你的变量名和路径格式来看,存储大概率是Azure Data Lake Storage,可直接用Azure官方Scala SDK读取文件:
import com.azure.storage.file.datalake.DataLakeFileClient import com.azure.storage.file.datalake.DataLakeFileSystemClient import com.azure.storage.file.datalake.DataLakeServiceClientBuilder // 初始化服务客户端 val serviceClient = new DataLakeServiceClientBuilder() .endpoint(s"https://${source_full_storage_account_name}.dfs.core.windows.net") .credential(/* 填入你的身份凭证,比如DefaultAzureCredential */) .buildClient() // 获取文件系统和文件客户端 val fileSystemClient: DataLakeFileSystemClient = serviceClient.getFileSystemClient("table1") val fileClient: DataLakeFileClient = fileSystemClient.getFileClient(s"/test/Tables/${table_name}.sql") // 读取文件内容 val upd_qry = new String(fileClient.readAllBytes())
方案2:将分布式存储挂载到本地
如果运行环境支持,可通过工具把ADLS存储挂载到本地文件系统(比如用Azure CLI的az storage fs mount命令),挂载后就能直接用Source.fromFile访问挂载后的本地路径。
方案3:基于Spark环境读取(大数据场景适配)
如果你的代码在Spark集群中运行,用Spark API读取更便捷:
import org.apache.spark.sql.SparkSession val spark = SparkSession.builder().getOrCreate() val upd_qry = spark.read.textFile(s"${dir}/${table_name}.sql").collect().mkString("\n")
内容的提问来源于stack exchange,提问作者user22067611
相关产品推荐
相关产品推荐

