如何批量从无表头文件提取指定列到Pandas DataFrame
Got it, let's break down how to solve this batch processing problem. Since your original CSV had a header "operation", but these split files don't have any headers, we just need to map that header to its column index (remember pandas uses 0-based indexing) and then process all files in bulk.
Step 1: Confirm the Target Column Index
First, figure out which column corresponds to your original "operation" column. For example:
- If "operation" was the first column in your original CSV, the index is
0 - If it was the third column, use
2
This is critical because we're working with headerless files now.
Step 2: Batch Processing Options
Option 1: Load All Data into Memory (Good for Smaller Files)
If your split files aren't too large, you can load all target columns into memory first, then combine them into a single file:
import pandas as pd import glob # Set your target column index here TARGET_COL_INDEX = 0 # Update this to match your actual column position # Get all files starting with "xxa" file_paths = glob.glob("xxa*") # Collect all target columns collected_data = [] for file in file_paths: # Read headerless CSV: header=None tells pandas first row is data, not headers df = pd.read_csv(file, header=None) # Extract the target column using iloc (for position-based indexing) target_col = df.iloc[:, TARGET_COL_INDEX] collected_data.append(target_col) # Combine all collected columns into one Series combined_series = pd.concat(collected_data, ignore_index=True) # Save to a new CSV with the original header combined_series.to_csv("extracted_operation_column.csv", header=["operation"], index=False)
Option 2: Stream Data to Output (Better for Large Files)
If you're dealing with large files that might eat up memory, process each file and append directly to the output file instead of loading everything at once:
import pandas as pd import glob TARGET_COL_INDEX = 0 file_paths = glob.glob("xxa*") # Write the header first with open("extracted_operation_column.csv", "w") as out_file: out_file.write("operation\n") # Loop through each file and append the target column for file in file_paths: try: # Only load the target column to save memory (usecols parameter) df = pd.read_csv(file, header=None, usecols=[TARGET_COL_INDEX]) # Append to output file, skip header since we already wrote it df.to_csv("extracted_operation_column.csv", mode="a", header=False, index=False) print(f"Processed: {file}") except Exception as e: print(f"Error processing {file}: {str(e)}")
Key Notes
- Encoding: If your files use a non-standard encoding (like
latin-1orutf-16), add theencodingparameter topd.read_csv()(e.g.,encoding="utf-8"). - Error Handling: The try-except block in Option 2 ensures a single bad file won't crash the entire batch process.
- File Matching:
glob.glob("xxa*")will match all files starting with "xxa" in your current working directory. If the files are in a subfolder, use something likeglob.glob("./path/to/files/xxa*").
内容的提问来源于stack exchange,提问作者Dileesh Dil

