如何在Databricks中使用Pandas的pd.read_excel读取/FileStore/tables/目录下的Excel文件及解决FileNotFoundError问题
I’ve run into this exact issue before, so let’s break down the steps to get your Excel file loaded into a pandas DataFrame successfully.
Step 1: Double-check the exact filename in DBFS
When you upload files via the Databricks UI, the platform often appends a random alphanumeric suffix to the filename (like abc_12345.xlsx) to avoid overwriting existing files. This is probably why your path isn’t working even though you see the file in the UI.
To confirm the actual filename, run this command in a notebook cell:
dbutils.fs.ls("/FileStore/tables/")
Look for your file in the output—note the full name including any suffixes. Then update your pandas read command with the correct path, e.g.:
import pandas as pd df = pd.read_excel("/dbfs/FileStore/tables/abc_12345.xlsx") display(df)
Step 2: If direct DBFS access still fails, copy to a local temp directory
In some cluster configurations, pandas might have trouble reading directly from the DBFS mount. A reliable workaround is to copy the file to Databricks' local temporary storage first, then read it:
# Copy the file from DBFS to local tmp directory dbutils.fs.cp("/FileStore/tables/abc_correct_name.xlsx", "file:/tmp/abc.xlsx") # Read from local tmp import pandas as pd df = pd.read_excel("/tmp/abc.xlsx") display(df)
Step 3: Read directly into memory (no local file copy)
If you prefer not to create a local copy, you can read the file content into a BytesIO object and pass it to pandas:
from io import BytesIO import pandas as pd # Read the file content as bytes file_content = dbutils.fs.cp("/FileStore/tables/abc_correct_name.xlsx", BytesIO()) # Reset the BytesIO pointer to the start file_content.seek(0) # Load into pandas DataFrame df = pd.read_excel(file_content) display(df)
All these methods are pure Python, so they should work within your organization's restrictions on Scala code.
内容的提问来源于stack exchange,提问作者Chaitanya Vardhan

