You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中校验命令行传入的CSV文件是否符合指定格式?

Got it, let's tackle this problem step by step. You need a Python program that takes a CSV path via command line, validates it against your specified format before any preprocessing, and throws clear errors if it's invalid. Here's a complete, practical solution:

Python CSV Format Validator with Command Line Input

1. Handle Command Line Input

First, we'll use Python's built-in argparse module to safely accept the CSV file path, plus basic checks for file existence and correct extension:

import argparse
import csv
from pathlib import Path

def parse_command_line_args():
    parser = argparse.ArgumentParser(description="Validate and process a CSV file with a specific format.")
    parser.add_argument("csv_path", type=str, help="Full path to your input CSV file")
    args = parser.parse_args()
    
    # Check if file exists and is a CSV
    csv_file = Path(args.csv_path)
    if not csv_file.exists():
        raise FileNotFoundError(f"File not found: '{csv_file}'")
    if csv_file.suffix.lower() != ".csv":
        raise ValueError("Input must be a CSV file (use .csv extension)")
    
    return csv_file

2. Core CSV Format Validation

Next, we'll write a function to validate every aspect of your required format:

  • Correct header columns (Sr.no, Codes, v1 to v300)
  • Valid integer values for Sr.no
  • Non-empty Codes field
  • Values in v1-v300 are either numbers or NA
def validate_csv_structure(csv_file):
    # Define the exact required columns
    required_columns = ["Sr.no", "Codes"] + [f"v{i}" for i in range(1, 301)]
    
    with open(csv_file, mode='r', newline='', encoding='utf-8') as f:
        reader = csv.DictReader(f)
        
        # Validate header matches exactly
        actual_columns = reader.fieldnames
        if actual_columns != required_columns:
            missing_cols = set(required_columns) - set(actual_columns)
            extra_cols = set(actual_columns) - set(required_columns)
            error_msg = "Invalid CSV header:\n"
            if missing_cols:
                error_msg += f"Missing columns: {', '.join(missing_cols)}\n"
            if extra_cols:
                error_msg += f"Unexpected columns: {', '.join(extra_cols)}\n"
            error_msg += f"Expected columns: {', '.join(required_columns[:5])} ... {required_columns[-1]}"
            raise ValueError(error_msg)
        
        # Validate each row's data
        for row_number, row in enumerate(reader, start=2):  # Row 1 is the header
            # Check Sr.no is a positive integer
            sr_no = row["Sr.no"].strip()
            if not sr_no.isdigit():
                raise ValueError(f"Row {row_number}: Sr.no '{sr_no}' must be an integer")
            
            # Check Codes isn't empty
            code = row["Codes"].strip()
            if not code:
                raise ValueError(f"Row {row_number}: Codes cannot be empty")
            
            # Check v1-v300 values are either NA or numeric
            for col in required_columns[2:]:
                value = row[col].strip()
                if value != "NA":
                    try:
                        # Try converting to float (works for integers too)
                        float(value)
                    except ValueError:
                        raise ValueError(f"Row {row_number}: Column '{col}' has invalid value '{value}' — must be a number or 'NA'")

3. Main Workflow

Put it all together with error handling to give users clear feedback:

def main():
    try:
        csv_file = parse_command_line_args()
        validate_csv_structure(csv_file)
        print(f"✅ CSV file '{csv_file}' is valid! Starting preprocessing...")
        # Add your preprocessing logic here
    except Exception as e:
        print(f"❌ Error: {str(e)}")
        exit(1)

if __name__ == "__main__":
    main()

Key Notes

  • Clear Error Messages: Every validation failure tells the user exactly what's wrong (missing columns, invalid value in a specific row/column) so they can fix their CSV easily.
  • Standard Libraries: No external dependencies — uses only Python's built-in argparse, csv, and pathlib modules.
  • Case Sensitivity: The code checks for exactly NA (uppercase). If your CSV might use lowercase na or other variants, you can modify the check to value.lower() == "na".

内容的提问来源于stack exchange,提问作者mjennet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:32:37