通过Widgets动态传参的Scala读取指定层级Parquet文件需求
Got it, let's break down how to implement this step by step—since you need to read Parquet files under the process/YYYY structure and pass month/day parameters dynamically using widgets, here's a practical approach:
First, Set Up Input Widgets
We'll use widgets (common in Spark environments like Databricks) to let users input or select the month and day. You can use text boxes for free input, or dropdowns to restrict values to valid options (like 01-12 for months):
// Create text widgets for month (MM) and day (DD) with default values dbutils.widgets.text("input_month", "01", "Enter Month (2-digit: 01-12)") dbutils.widgets.text("input_day", "01", "Enter Day (2-digit: 01-31)") // Optional: Switch to dropdowns for predefined valid values // dbutils.widgets.dropdown("input_month", "01", (1 to 12).map(n => f"$n%02d").toSeq, "Select Month") // dbutils.widgets.dropdown("input_day", "01", (1 to 31).map(n => f"$n%02d").toSeq, "Select Day")
Next, Grab and Validate Widget Values
It's a good idea to validate the input to make sure it's in the correct 2-digit format—this avoids errors when building the file path:
val month = dbutils.widgets.get("input_month") val day = dbutils.widgets.get("input_day") // Basic validation for 2-digit month/day if (!month.matches("^0[1-9]|1[0-2]$")) { throw new IllegalArgumentException("Oops! Month needs to be a 2-digit value between 01 and 12.") } if (!day.matches("^0[1-9]|[12][0-9]|3[01]$")) { throw new IllegalArgumentException("Oops! Day needs to be a 2-digit value between 01 and 31.") }
Build the Path and Read Parquet Files
Now construct the target path using your fixed year (replace 2024 with your actual year, or add another widget for year if you need that too) and the dynamic month/day values. Then read the Parquet data:
// Replace 2024 with your target year (or add a widget for year flexibility) val baseYearPath = "/mnt/pnt/process/2024" val targetPath = s"$baseYearPath/$month/$day" // Read the Parquet files from the specific month/day folder val readDf = spark.read.format("parquet").load(targetPath) // Alternative: If your data is partitioned by month/day, read the whole year and filter // val readDf = spark.read.format("parquet").load(baseYearPath) // .filter(s"month = '$month' AND day = '$day'")
Quick Check to Confirm
Add a quick verification step to make sure the data loaded properly:
// Show a sample of the data and print the schema readDf.show(5) readDf.printSchema()
A Few Extra Tips:
- If your Parquet data is partitioned by
monthandday, filtering after reading the entire year path can be more efficient (Spark can prune partitions automatically). - Use dropdown widgets instead of text boxes if you want to prevent invalid inputs entirely.
- Double-check that your Spark cluster has access to the
/mnt/pnt/processpath (especially if you're using Databricks mounts).
内容的提问来源于stack exchange,提问作者Pradyot Mohanty

