如何实现从多个CSV文件中提取第N行并写入新文件?
Hey Mike, let's get that multi-file row extraction sorted out! It sounds like your current setup is only pulling the target line from the first CSV in your folder—totally a common pitfall when looping through files. Here are two straightforward solutions to make this work for every file:
Python Solution (Flexible & Customizable)
Since you already have a working single-file method, let's expand that into a loop that handles all CSVs. This example extracts the 5th line (note: if your CSV has a header, adjust the index accordingly—we use 4 here because Python uses 0-based indexing):
import os # Configuration target_line_number = 5 # The line you want to extract (1-based) input_folder = "./your_csv_folder" # Path to your CSV folder output_file = "extracted_lines.csv" # Convert to 0-based index target_index = target_line_number - 1 # Clear output file first (optional, for fresh results) with open(output_file, "w") as out_f: pass # Loop through all CSV files in the folder for filename in os.listdir(input_folder): if filename.endswith(".csv"): file_path = os.path.join(input_folder, filename) try: with open(file_path, "r") as in_f: lines = in_f.readlines() # Check if the file has enough lines to avoid errors if len(lines) > target_index: target_line = lines[target_index].strip() # Append to output file (use "a" to avoid overwriting previous entries) with open(output_file, "a") as out_f: # Optional: Add filename as a prefix to track which line came from where out_f.write(f"{filename},{target_line}\n") else: print(f"Warning: {filename} has fewer than {target_line_number} lines") except Exception as e: print(f"Error processing {filename}: {str(e)}")
Why this works:
- We use
os.listdir()to fetch all files in your folder, then filter for.csvextensions to focus only on relevant files. - Opening the output file in append mode (
"a") ensures we don't overwrite previous entries—this is likely the fix for your original issue where only the first file's line was saved! - Basic error handling helps skip problematic files and alerts you to CSVs that are too short.
Command-Line Solution (Quick & No Code)
If you prefer terminal commands (great for Linux/macOS, or Windows with WSL), you can use awk and a simple loop to get the job done in one go. This extracts the 5th line from every CSV and appends it to extracted_lines.csv:
# For Linux/macOS/WSL for file in ./your_csv_folder/*.csv; do # Print the 5th line of each file, with the filename added for context awk -v fname="$file" 'NR==5 {print fname "," $0}' "$file" >> extracted_lines.csv done
Quick Checks for Your Original Script
If your existing code only works for the first file, verify these two common issues:
- Output file mode: If you're using
"w"instead of"a"when writing to the output file, it will overwrite the file every loop iteration, leaving only the first file's line. - Loop logic: Double-check that your code is actually iterating over all CSV files—make sure you didn't accidentally add a
breakstatement, or that your file list isn't truncated after the first entry.
内容的提问来源于stack exchange,提问作者MikeeeeeeeyC

