You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Scala读取Parquet文件时如何保留包含空Parquet文件的目录以满足日志需求

Solution for Retaining Empty Parquet Directory Metadata During Read

Got it, here are a couple of practical approaches to handle this scenario—you’ll be able to read the valid data from nm4 while capturing all critical metadata about nm3 (like that it exists but contains empty parquet files) for your logging needs.

Approach 1: Explicit Directory Iteration with Logging (Pandas/PyArrow)

The most straightforward way is to loop through each subdirectory individually, read its parquet files, and log the status of each. This ensures you don’t miss any directory info, even if it’s empty.

import os
import pandas as pd
import logging

# Set up logging to track directory processing
logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

# Define your base path containing nm3 and nm4
base_path = "fld/fld1/fld2"

# Get all subdirectories under the base path
subdirectories = [d for d in os.listdir(base_path) if os.path.isdir(os.path.join(base_path, d))]

combined_data = pd.DataFrame()
directory_logs = []

for subdir in subdirectories:
    dir_full_path = os.path.join(base_path, subdir)
    try:
        # Read all parquet files in the current subdirectory
        df = pd.read_parquet(dir_full_path)
        row_count = len(df)
        
        # Log the directory's status
        logger.info(f"Processed directory '{subdir}': {row_count} rows found")
        directory_logs.append({
            "directory_name": subdir,
            "row_count": row_count,
            "status": "success"
        })
        
        # Append non-empty data to your combined dataset
        if row_count > 0:
            combined_data = pd.concat([combined_data, df], ignore_index=True)
    
    except Exception as e:
        # Log any errors encountered (e.g., corrupted files)
        logger.error(f"Failed to process directory '{subdir}': {str(e)}")
        directory_logs.append({
            "directory_name": subdir,
            "row_count": 0,
            "status": "error",
            "error_message": str(e)
        })

# Your final data (only nm4's content) and logs (includes nm3's empty status)
print(f"Combined data shape: {combined_data.shape}")
print("Directory processing logs:")
for log in directory_logs:
    print(log)

Why this works:

  • You explicitly process each subdirectory, so nm3 is never overlooked.
  • Even if nm3’s parquet files are empty, you log that it had 0 rows, which is exactly the info you need for auditing.
  • The combined dataset only includes valid data from nm4, keeping your analysis clean.

Approach 2: Granular File Inspection with PyArrow

If you want more detailed metadata (like how many empty parquet files are in nm3), use PyArrow’s filesystem tools to inspect files before reading them:

import pyarrow.parquet as pq
import pyarrow.fs as fs
import os
import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

base_path = "fld/fld1/fld2"
local_filesystem = fs.LocalFileSystem()

# Get all subdirectories
subdirs = [entry.path for entry in local_filesystem.get_file_info(base_path) if entry.is_dir]

for subdir in subdirs:
    subdir_name = os.path.basename(subdir)
    # List all parquet files in the subdirectory
    parquet_files = [f.path for f in local_filesystem.get_file_info(subdir) if f.path.endswith(".parquet")]
    
    if not parquet_files:
        logger.info(f"Directory '{subdir_name}' has no parquet files")
        continue
    
    # Count empty files (size = 0 bytes)
    empty_file_count = sum(1 for f in local_filesystem.get_file_info(parquet_files) if f.size == 0)
    total_file_count = len(parquet_files)
    
    logger.info(f"Directory '{subdir_name}': {total_file_count} parquet files, {empty_file_count} empty")
    
    # Read data only if there are non-empty files
    if empty_file_count != total_file_count:
        table = pq.read_table(subdir)
        df = table.to_pandas()
        logger.info(f"Loaded {len(df)} rows from '{subdir_name}'")
        # Process df as needed
    else:
        logger.info(f"No valid data to load from '{subdir_name}'")

Why this works:

  • You get deeper insights into the state of nm3 (e.g., "3 empty parquet files") instead of just "0 rows".
  • Avoids reading entirely empty directories unnecessarily, saving a bit of processing time.

Key Takeaway

The main issue with reading the parent directory directly is that while it will include empty rows from nm3, you lose the context of which directory those empty rows came from. By iterating over each subdirectory explicitly, you retain full visibility into each path’s status—perfect for your logging requirements.

内容的提问来源于stack exchange,提问作者justanothertekguy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:47:51