You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:文件缺失时跳过对应处理块的代码修改需求

Fixing Pandas Code to Skip Missing Files & Reduce Repetition

Looks like you're dealing with two main pain points here: your script crashes when a file is missing, and you've got tons of repetitive code that's error-prone and hard to maintain. Let's fix both issues with a cleaner, more robust approach.

First, let's break down what's causing problems in your original code:

  • No error handling for missing files: pd.read_csv throws a FileNotFoundError if the file doesn't exist, stopping your entire workflow.
  • Repetitive logic: You're copying the exact same processing steps for every file—this leads to typos (like the JPH_regressions mistake in your 8th block) and makes updates a hassle.
  • Bad variable naming: Starting variable names with numbers (e.g., 1regressions) is against Python best practices and can cause unexpected issues.

Here's the improved version that skips missing files and eliminates redundancy:

import pandas as pd
import os

# Base path for your files
base_path = "/Users/xyz"
# List of all regression files we want to process
file_list = [
    "11regressions.csv",
    "12regressions.csv",
    "13regressions.csv",
    "14regressions.csv",
    "15regressions.csv",
    "16regressions.csv",
    "17regressions.csv",
    "18regressions.csv"
]
# Define the columns we want to keep in the final output
target_columns = [
    'filename', 'metric_type', 'content_format',
    'tp50_pct', 'tp90_pct', 'mean_pct',
    'iterations', 'perf_run_id', 'baseline_perf_run_id'
]

# Empty list to store processed DataFrames
processed_data = []

for filename in file_list:
    full_path = os.path.join(base_path, filename)
    
    # Skip if the file doesn't exist
    if not os.path.exists(full_path):
        print(f"Skipping missing file: {filename}")
        continue
    
    try:
        # Read and process the CSV
        df = pd.read_csv(full_path)
        # Filter rows where tp50 or tp90 exceeds 0.15
        df = df[(df['tp50_pct'] > 0.15) | (df['tp90_pct'] > 0.15)]
        # Add filename identifier
        df['filename'] = filename
        # Keep only the columns we care about
        df = df[target_columns]
        # Add to our processed list
        processed_data.append(df)
    except Exception as e:
        print(f"Error processing {filename}: {str(e)}")
        continue

# Combine all valid data and save
if processed_data:
    result = pd.concat(processed_data, ignore_index=True)
    # Reset index to start at 1 instead of 0
    result.index = range(1, len(result) + 1)
    result.to_csv(os.path.join(base_path, "test.csv"), index=False)
else:
    print("No valid data to save—all files were missing or corrupted.")

Key Improvements:

  • Graceful file skipping: Uses os.path.exists() to check for missing files and logs a message instead of crashing.
  • Error resilience: The try-except block catches other issues (like corrupted CSVs) so your script keeps running.
  • No repetition: All processing logic lives in one loop—update the filter or columns once, and it applies to all files.
  • Better practices:
    • os.path.join() builds file paths safely (avoids missing slashes or cross-platform issues).
    • pd.concat() replaces the deprecated append() method for combining DataFrames.
    • Clear, maintainable lists for files and columns make future edits easy.
    • Checks if there's valid data before saving to avoid empty output files.

内容的提问来源于stack exchange,提问作者Vasantham S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:47:48