Python处理CSV文件:删除无用行列并保存(无NumPy/Pandas)
I'll break down solutions for both your requirements using basic Python file handling and string operations—no external libraries needed.
Requirement 1: Remove Specific Column and Fixed Rows
For your exact case where you need to delete the 2nd column, first 3 rows, and last 3 rows, here's a straightforward implementation:
# Clean CSV per Requirement 1 input_path = "input.csv" output_path = "cleaned_req1.csv" # Read all lines, filter out empty ones with open(input_path, 'r') as infile: lines = [line.strip() for line in infile if line.strip()] # Keep lines excluding first 3 and last 3 filtered_lines = lines[3:-3] # Process each line to remove the 2nd column (index 1) processed = [] for line in filtered_lines: cols = line.split(',') # Keep all columns except the second one cleaned_cols = cols[:1] + cols[2:] processed.append(','.join(cleaned_cols)) # Write the result to a new file with open(output_path, 'w') as outfile: outfile.write('\n'.join(processed))
How it works:
- Reads all lines from the input CSV, stripping whitespace and skipping empty lines.
- Uses list slicing to exclude the first 3 and last 3 rows.
- Splits each line into columns, removes the second column (index 1), then joins the columns back into a string.
- Writes the cleaned lines to a new CSV file.
Requirement 2: Dynamic Handling of Useless Rows/Columns
If you have more than 10 useless rows (only at the start/end) where valid rows start with a_, and multiple useless columns, we can adjust the code to dynamically identify valid rows instead of using fixed indices.
Solution Code:
# Clean CSV per Requirement 2 input_path = "input.csv" output_path = "cleaned_req2.csv" # Define which columns to remove (use indices, e.g., {1} for 2nd column) useless_col_indices = {1} # Read all lines and filter out empty ones with open(input_path, 'r') as infile: lines = [line.strip() for line in infile if line.strip()] # Find indices of all valid rows (first column starts with 'a_') valid_row_positions = [] for idx, line in enumerate(lines): first_col = line.split(',')[0].strip() if first_col.startswith('a_'): valid_row_positions.append(idx) if not valid_row_positions: print("Error: No valid rows found in the CSV.") else: # Extract only valid lines from first to last valid row valid_lines = lines[valid_row_positions[0]: valid_row_positions[-1] + 1] # Process lines to remove useless columns processed = [] for line in valid_lines: cols = line.split(',') cleaned_cols = [col for idx, col in enumerate(cols) if idx not in useless_col_indices] processed.append(','.join(cleaned_cols)) # Write the cleaned data with open(output_path, 'w') as outfile: outfile.write('\n'.join(processed))
How it works:
- Identifies all rows where the first column starts with
a_, regardless of how many useless rows are at the start or end. - Uses a set of indices to specify which columns to remove (you can add multiple indices like
{1, 3}if needed). - Extracts only the valid rows and removes the specified columns before writing to the output file.
Handling Large CSV Files
If your CSV is too big to load all lines into memory, you can process it in two passes to avoid high memory usage:
# Efficient solution for large CSV files (Requirement 2) input_path = "large_input.csv" output_path = "cleaned_large.csv" useless_col_indices = {1} # First pass: find the last valid line last_valid_line = None with open(input_path, 'r') as infile: for line in infile: stripped_line = line.strip() if not stripped_line: continue first_col = stripped_line.split(',')[0].strip() if first_col.startswith('a_'): last_valid_line = stripped_line if not last_valid_line: print("Error: No valid rows found.") else: # Second pass: write valid lines from first valid to last valid with open(input_path, 'r') as infile, open(output_path, 'w') as outfile: writing = False for line in infile: stripped_line = line.strip() if not stripped_line: continue first_col = stripped_line.split(',')[0].strip() # Start writing when we hit the first valid row if first_col.startswith('a_'): writing = True # Clean the line by removing useless columns cols = stripped_line.split(',') cleaned_cols = [col for idx, col in enumerate(cols) if idx not in useless_col_indices] cleaned_line = ','.join(cleaned_cols) if writing: outfile.write(cleaned_line + '\n') # Stop writing after we process the last valid line if cleaned_line == ','.join([col for idx, col in enumerate(last_valid_line.split(',')) if idx not in useless_col_indices]): break
This approach reads the file twice (once to find the last valid row, once to write the valid range) which uses minimal memory even for huge files.
内容的提问来源于stack exchange,提问作者kts

