Spark Scala读取Parquet文件时如何保留包含空Parquet文件的目录以满足日志需求
Got it, here are a couple of practical approaches to handle this scenario—you’ll be able to read the valid data from nm4 while capturing all critical metadata about nm3 (like that it exists but contains empty parquet files) for your logging needs.
Approach 1: Explicit Directory Iteration with Logging (Pandas/PyArrow)
The most straightforward way is to loop through each subdirectory individually, read its parquet files, and log the status of each. This ensures you don’t miss any directory info, even if it’s empty.
import os import pandas as pd import logging # Set up logging to track directory processing logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) # Define your base path containing nm3 and nm4 base_path = "fld/fld1/fld2" # Get all subdirectories under the base path subdirectories = [d for d in os.listdir(base_path) if os.path.isdir(os.path.join(base_path, d))] combined_data = pd.DataFrame() directory_logs = [] for subdir in subdirectories: dir_full_path = os.path.join(base_path, subdir) try: # Read all parquet files in the current subdirectory df = pd.read_parquet(dir_full_path) row_count = len(df) # Log the directory's status logger.info(f"Processed directory '{subdir}': {row_count} rows found") directory_logs.append({ "directory_name": subdir, "row_count": row_count, "status": "success" }) # Append non-empty data to your combined dataset if row_count > 0: combined_data = pd.concat([combined_data, df], ignore_index=True) except Exception as e: # Log any errors encountered (e.g., corrupted files) logger.error(f"Failed to process directory '{subdir}': {str(e)}") directory_logs.append({ "directory_name": subdir, "row_count": 0, "status": "error", "error_message": str(e) }) # Your final data (only nm4's content) and logs (includes nm3's empty status) print(f"Combined data shape: {combined_data.shape}") print("Directory processing logs:") for log in directory_logs: print(log)
Why this works:
- You explicitly process each subdirectory, so
nm3is never overlooked. - Even if
nm3’s parquet files are empty, you log that it had 0 rows, which is exactly the info you need for auditing. - The combined dataset only includes valid data from
nm4, keeping your analysis clean.
Approach 2: Granular File Inspection with PyArrow
If you want more detailed metadata (like how many empty parquet files are in nm3), use PyArrow’s filesystem tools to inspect files before reading them:
import pyarrow.parquet as pq import pyarrow.fs as fs import os import logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger(__name__) base_path = "fld/fld1/fld2" local_filesystem = fs.LocalFileSystem() # Get all subdirectories subdirs = [entry.path for entry in local_filesystem.get_file_info(base_path) if entry.is_dir] for subdir in subdirs: subdir_name = os.path.basename(subdir) # List all parquet files in the subdirectory parquet_files = [f.path for f in local_filesystem.get_file_info(subdir) if f.path.endswith(".parquet")] if not parquet_files: logger.info(f"Directory '{subdir_name}' has no parquet files") continue # Count empty files (size = 0 bytes) empty_file_count = sum(1 for f in local_filesystem.get_file_info(parquet_files) if f.size == 0) total_file_count = len(parquet_files) logger.info(f"Directory '{subdir_name}': {total_file_count} parquet files, {empty_file_count} empty") # Read data only if there are non-empty files if empty_file_count != total_file_count: table = pq.read_table(subdir) df = table.to_pandas() logger.info(f"Loaded {len(df)} rows from '{subdir_name}'") # Process df as needed else: logger.info(f"No valid data to load from '{subdir_name}'")
Why this works:
- You get deeper insights into the state of
nm3(e.g., "3 empty parquet files") instead of just "0 rows". - Avoids reading entirely empty directories unnecessarily, saving a bit of processing time.
Key Takeaway
The main issue with reading the parent directory directly is that while it will include empty rows from nm3, you lose the context of which directory those empty rows came from. By iterating over each subdirectory explicitly, you retain full visibility into each path’s status—perfect for your logging requirements.
内容的提问来源于stack exchange,提问作者justanothertekguy

