Azure Databricks中Selenium Python自动化文件下载路径及存储咨询
Hey there! Let's tackle your questions about exporting Excel files via Selenium in Azure Databricks, including finding the download path and sending files directly to storage like Azure Blob.
1. Where does the Excel file download in Azure Databricks?
When running Selenium Chrome on an Azure Databricks cluster, the browser's default download path is tied to the cluster node's local filesystem—usually something like /tmp/<your-username>/downloads. But here's the catch: this local path is temporary—if the cluster restarts or the node is recycled, the file will be gone.
A better approach is to explicitly set a download directory pointing to DBFS (Databricks File System). DBFS is a distributed, persistent storage layer accessible across all cluster nodes, and it integrates seamlessly with other Azure services. Here's how to configure this in your code:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import time import os # Configure Chrome to download to a DBFS path chrome_options = Options() # Use a DBFS path (prefix with /dbfs/ to access it from the node's filesystem) download_directory = "/dbfs/tmp/excel_exports" # Set Chrome preferences for downloads chrome_options.add_experimental_option("prefs", { "download.default_directory": download_directory, "download.prompt_for_download": False, # Disable download prompt "download.directory_upgrade": True, "safebrowsing.enabled": True # Skip safety checks for trusted downloads }) # Initialize the Chrome driver with these options driver = webdriver.Chrome(options=chrome_options) # Your existing export logic driver.get("your-target-url") exportToExcel = driver.find_element_by_xpath('//*[@id="excelReport"]') exportToExcel.click() # Instead of fixed sleep, wait for the file to appear (more reliable) downloaded_file = None timeout = 30 # Wait up to 30 seconds start_time = time.time() while time.time() - start_time < timeout: # List files in the download directory files = os.listdir(download_directory) for file in files: if file.endswith(".xlsx") and not file.endswith(".crdownload"): # Skip incomplete downloads downloaded_file = os.path.join(download_directory, file) break if downloaded_file: break time.sleep(1) if downloaded_file: print(f"Excel file downloaded to: {downloaded_file}") else: print("Download timed out.") driver.quit()
2. Can I download directly to Azure Blob Storage or other specified storage?
Selenium's Chrome driver can't download directly to Azure Blob Storage—browser downloads are restricted to local/filesystem paths. But you can easily move the downloaded file from DBFS to Blob Storage using Databricks built-in tools or Azure SDKs.
Option 1: Use dbutils.fs (simplest for Databricks)
If your Azure Blob Storage is already mounted to DBFS, or you have access via a storage account key/SAS token, use dbutils.fs.cp to copy the file:
# Define paths source_dbfs_path = "dbfs:/tmp/excel_exports/report.xlsx" # Match your download directory target_blob_path = "wasbs://<your-container>@<your-storage-account>.blob.core.windows.net/reports/report.xlsx" # Copy the file to Blob Storage dbutils.fs.cp(source_dbfs_path, target_blob_path) print(f"File copied to Blob Storage: {target_blob_path}")
Option 2: Use Azure Storage SDK
If you need more control (like setting metadata, handling large files), use the azure-storage-blob library. First install it on your cluster (via Libraries tab or %pip install azure-storage-blob), then run:
from azure.storage.blob import BlobServiceClient # Initialize Blob Service Client connection_string = "your-storage-account-connection-string" blob_service_client = BlobServiceClient.from_connection_string(connection_string) # Get blob container client container_client = blob_service_client.get_container_client("<your-container>") # Upload the file from DBFS to Blob with open(downloaded_file, "rb") as data: container_client.upload_blob(name="reports/report.xlsx", data=data, overwrite=True) print("File uploaded to Blob Storage successfully.")
Key Notes to Remember
- Cluster Setup: Ensure your Databricks cluster has Chrome and ChromeDriver installed. You can use an init script to install them automatically when the cluster starts.
- Permissions: Make sure your cluster has access to the target Blob Storage (via Service Principal, SAS token, or storage account key).
- Avoid
time.sleep(): Using a loop to check for the downloaded file is more reliable than fixed sleep times, especially in variable cluster environments.
内容的提问来源于stack exchange,提问作者Karthick Kumar

