Python Pandas:文件缺失时跳过对应处理块的代码修改需求
Fixing Pandas Code to Skip Missing Files & Reduce Repetition
Looks like you're dealing with two main pain points here: your script crashes when a file is missing, and you've got tons of repetitive code that's error-prone and hard to maintain. Let's fix both issues with a cleaner, more robust approach.
First, let's break down what's causing problems in your original code:
- No error handling for missing files:
pd.read_csvthrows aFileNotFoundErrorif the file doesn't exist, stopping your entire workflow. - Repetitive logic: You're copying the exact same processing steps for every file—this leads to typos (like the
JPH_regressionsmistake in your 8th block) and makes updates a hassle. - Bad variable naming: Starting variable names with numbers (e.g.,
1regressions) is against Python best practices and can cause unexpected issues.
Here's the improved version that skips missing files and eliminates redundancy:
import pandas as pd import os # Base path for your files base_path = "/Users/xyz" # List of all regression files we want to process file_list = [ "11regressions.csv", "12regressions.csv", "13regressions.csv", "14regressions.csv", "15regressions.csv", "16regressions.csv", "17regressions.csv", "18regressions.csv" ] # Define the columns we want to keep in the final output target_columns = [ 'filename', 'metric_type', 'content_format', 'tp50_pct', 'tp90_pct', 'mean_pct', 'iterations', 'perf_run_id', 'baseline_perf_run_id' ] # Empty list to store processed DataFrames processed_data = [] for filename in file_list: full_path = os.path.join(base_path, filename) # Skip if the file doesn't exist if not os.path.exists(full_path): print(f"Skipping missing file: {filename}") continue try: # Read and process the CSV df = pd.read_csv(full_path) # Filter rows where tp50 or tp90 exceeds 0.15 df = df[(df['tp50_pct'] > 0.15) | (df['tp90_pct'] > 0.15)] # Add filename identifier df['filename'] = filename # Keep only the columns we care about df = df[target_columns] # Add to our processed list processed_data.append(df) except Exception as e: print(f"Error processing {filename}: {str(e)}") continue # Combine all valid data and save if processed_data: result = pd.concat(processed_data, ignore_index=True) # Reset index to start at 1 instead of 0 result.index = range(1, len(result) + 1) result.to_csv(os.path.join(base_path, "test.csv"), index=False) else: print("No valid data to save—all files were missing or corrupted.")
Key Improvements:
- Graceful file skipping: Uses
os.path.exists()to check for missing files and logs a message instead of crashing. - Error resilience: The
try-exceptblock catches other issues (like corrupted CSVs) so your script keeps running. - No repetition: All processing logic lives in one loop—update the filter or columns once, and it applies to all files.
- Better practices:
os.path.join()builds file paths safely (avoids missing slashes or cross-platform issues).pd.concat()replaces the deprecatedappend()method for combining DataFrames.- Clear, maintainable lists for files and columns make future edits easy.
- Checks if there's valid data before saving to avoid empty output files.
内容的提问来源于stack exchange,提问作者Vasantham S
相关产品推荐
相关产品推荐

