如何在Python中校验命令行传入的CSV文件是否符合指定格式?
Got it, let's tackle this problem step by step. You need a Python program that takes a CSV path via command line, validates it against your specified format before any preprocessing, and throws clear errors if it's invalid. Here's a complete, practical solution:
Python CSV Format Validator with Command Line Input
1. Handle Command Line Input
First, we'll use Python's built-in argparse module to safely accept the CSV file path, plus basic checks for file existence and correct extension:
import argparse import csv from pathlib import Path def parse_command_line_args(): parser = argparse.ArgumentParser(description="Validate and process a CSV file with a specific format.") parser.add_argument("csv_path", type=str, help="Full path to your input CSV file") args = parser.parse_args() # Check if file exists and is a CSV csv_file = Path(args.csv_path) if not csv_file.exists(): raise FileNotFoundError(f"File not found: '{csv_file}'") if csv_file.suffix.lower() != ".csv": raise ValueError("Input must be a CSV file (use .csv extension)") return csv_file
2. Core CSV Format Validation
Next, we'll write a function to validate every aspect of your required format:
- Correct header columns (Sr.no, Codes, v1 to v300)
- Valid integer values for
Sr.no - Non-empty
Codesfield - Values in v1-v300 are either numbers or
NA
def validate_csv_structure(csv_file): # Define the exact required columns required_columns = ["Sr.no", "Codes"] + [f"v{i}" for i in range(1, 301)] with open(csv_file, mode='r', newline='', encoding='utf-8') as f: reader = csv.DictReader(f) # Validate header matches exactly actual_columns = reader.fieldnames if actual_columns != required_columns: missing_cols = set(required_columns) - set(actual_columns) extra_cols = set(actual_columns) - set(required_columns) error_msg = "Invalid CSV header:\n" if missing_cols: error_msg += f"Missing columns: {', '.join(missing_cols)}\n" if extra_cols: error_msg += f"Unexpected columns: {', '.join(extra_cols)}\n" error_msg += f"Expected columns: {', '.join(required_columns[:5])} ... {required_columns[-1]}" raise ValueError(error_msg) # Validate each row's data for row_number, row in enumerate(reader, start=2): # Row 1 is the header # Check Sr.no is a positive integer sr_no = row["Sr.no"].strip() if not sr_no.isdigit(): raise ValueError(f"Row {row_number}: Sr.no '{sr_no}' must be an integer") # Check Codes isn't empty code = row["Codes"].strip() if not code: raise ValueError(f"Row {row_number}: Codes cannot be empty") # Check v1-v300 values are either NA or numeric for col in required_columns[2:]: value = row[col].strip() if value != "NA": try: # Try converting to float (works for integers too) float(value) except ValueError: raise ValueError(f"Row {row_number}: Column '{col}' has invalid value '{value}' — must be a number or 'NA'")
3. Main Workflow
Put it all together with error handling to give users clear feedback:
def main(): try: csv_file = parse_command_line_args() validate_csv_structure(csv_file) print(f"✅ CSV file '{csv_file}' is valid! Starting preprocessing...") # Add your preprocessing logic here except Exception as e: print(f"❌ Error: {str(e)}") exit(1) if __name__ == "__main__": main()
Key Notes
- Clear Error Messages: Every validation failure tells the user exactly what's wrong (missing columns, invalid value in a specific row/column) so they can fix their CSV easily.
- Standard Libraries: No external dependencies — uses only Python's built-in
argparse,csv, andpathlibmodules. - Case Sensitivity: The code checks for exactly
NA(uppercase). If your CSV might use lowercasenaor other variants, you can modify the check tovalue.lower() == "na".
内容的提问来源于stack exchange,提问作者mjennet
相关产品推荐
相关产品推荐

