如何使用PySpark从Azure Blob Storage读取WAV音频数据
在Databricks中用PySpark读取Azure Blob Storage音频数据的方案
我在Databricks环境下尝试用PySpark读取Azure Blob Storage中的音频数据时遇到了问题,目前已实现将音频数据读取为原始二进制文件的方案,但需要说明的是,PySpark并非为处理音频文件设计,因此该方案并非最优解。
配置并读取音频文件
# Settings storage_account_name = "storage_account_name" blob_name = "blob_name" storage_account_access_key = "key" file_name = "audio.wav" file_type = "binaryFile" # Spark configuration and file location spark.conf.set( "fs.azure.account.key."+storage_account_name+".blob.core.windows.net", storage_account_access_key) file_location = "wasbs://" + blob_name + "@" + storage_account_name + ".blob.core.windows.net/" + file_name df = spark.read.format(file_type).load(file_location)
读取并播放音频数据
# Convert spark dataframe to pandas dataframe df = df.toPandas() # Sound player from IPython.display import Audio, display display(Audio(df['content'].iloc[0]))
执行上述代码后,可在Notebook中看到音频播放器界面:
如果这个方案对你有帮助,欢迎点赞支持!
内容的提问来源于Stack Exchange,提问作者Vojtech Stas
相关产品推荐
相关产品推荐

