You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Databricks中Selenium Python自动化文件下载路径及存储咨询

Selenium Chrome Automation in Azure Databricks: Handling Excel Exports

Hey there! Let's tackle your questions about exporting Excel files via Selenium in Azure Databricks, including finding the download path and sending files directly to storage like Azure Blob.

1. Where does the Excel file download in Azure Databricks?

When running Selenium Chrome on an Azure Databricks cluster, the browser's default download path is tied to the cluster node's local filesystem—usually something like /tmp/<your-username>/downloads. But here's the catch: this local path is temporary—if the cluster restarts or the node is recycled, the file will be gone.

A better approach is to explicitly set a download directory pointing to DBFS (Databricks File System). DBFS is a distributed, persistent storage layer accessible across all cluster nodes, and it integrates seamlessly with other Azure services. Here's how to configure this in your code:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
import os

# Configure Chrome to download to a DBFS path
chrome_options = Options()
# Use a DBFS path (prefix with /dbfs/ to access it from the node's filesystem)
download_directory = "/dbfs/tmp/excel_exports"

# Set Chrome preferences for downloads
chrome_options.add_experimental_option("prefs", {
    "download.default_directory": download_directory,
    "download.prompt_for_download": False,  # Disable download prompt
    "download.directory_upgrade": True,
    "safebrowsing.enabled": True  # Skip safety checks for trusted downloads
})

# Initialize the Chrome driver with these options
driver = webdriver.Chrome(options=chrome_options)

# Your existing export logic
driver.get("your-target-url")
exportToExcel = driver.find_element_by_xpath('//*[@id="excelReport"]')
exportToExcel.click()

# Instead of fixed sleep, wait for the file to appear (more reliable)
downloaded_file = None
timeout = 30  # Wait up to 30 seconds
start_time = time.time()

while time.time() - start_time < timeout:
    # List files in the download directory
    files = os.listdir(download_directory)
    for file in files:
        if file.endswith(".xlsx") and not file.endswith(".crdownload"):  # Skip incomplete downloads
            downloaded_file = os.path.join(download_directory, file)
            break
    if downloaded_file:
        break
    time.sleep(1)

if downloaded_file:
    print(f"Excel file downloaded to: {downloaded_file}")
else:
    print("Download timed out.")

driver.quit()

2. Can I download directly to Azure Blob Storage or other specified storage?

Selenium's Chrome driver can't download directly to Azure Blob Storage—browser downloads are restricted to local/filesystem paths. But you can easily move the downloaded file from DBFS to Blob Storage using Databricks built-in tools or Azure SDKs.

Option 1: Use dbutils.fs (simplest for Databricks)

If your Azure Blob Storage is already mounted to DBFS, or you have access via a storage account key/SAS token, use dbutils.fs.cp to copy the file:

# Define paths
source_dbfs_path = "dbfs:/tmp/excel_exports/report.xlsx"  # Match your download directory
target_blob_path = "wasbs://<your-container>@<your-storage-account>.blob.core.windows.net/reports/report.xlsx"

# Copy the file to Blob Storage
dbutils.fs.cp(source_dbfs_path, target_blob_path)
print(f"File copied to Blob Storage: {target_blob_path}")

Option 2: Use Azure Storage SDK

If you need more control (like setting metadata, handling large files), use the azure-storage-blob library. First install it on your cluster (via Libraries tab or %pip install azure-storage-blob), then run:

from azure.storage.blob import BlobServiceClient

# Initialize Blob Service Client
connection_string = "your-storage-account-connection-string"
blob_service_client = BlobServiceClient.from_connection_string(connection_string)

# Get blob container client
container_client = blob_service_client.get_container_client("<your-container>")

# Upload the file from DBFS to Blob
with open(downloaded_file, "rb") as data:
    container_client.upload_blob(name="reports/report.xlsx", data=data, overwrite=True)

print("File uploaded to Blob Storage successfully.")

Key Notes to Remember

  • Cluster Setup: Ensure your Databricks cluster has Chrome and ChromeDriver installed. You can use an init script to install them automatically when the cluster starts.
  • Permissions: Make sure your cluster has access to the target Blob Storage (via Service Principal, SAS token, or storage account key).
  • Avoid time.sleep(): Using a loop to check for the downloaded file is more reliable than fixed sleep times, especially in variable cluster environments.

内容的提问来源于stack exchange,提问作者Karthick Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 12:57:53