如何用Python将输入文本文件的指定列分别保存到独立输出文件
Solution: Python Script to Extract Specific Columns to Separate Output Files
Got it, let's build exactly what you need step by step. Below is a complete septc.py script that handles your column extraction requirement, with explanations to help you follow along.
import argparse from contextlib import ExitStack def main(): # Set up command-line argument parser parser = argparse.ArgumentParser(description='Pull specific columns from a text file into separate output files.') parser.add_argument('--infile', required=True, help='Path to your input text file') parser.add_argument('--column', required=True, nargs='+', type=int, help='0-based indices of columns to extract (e.g., 0 2 3)') parser.add_argument('--outfile', required=True, nargs='+', help='Output file paths (one for each column, in order)') args = parser.parse_args() # First, make sure column count matches output file count if len(args.column) != len(args.outfile): print("Error: Number of columns must match number of output files!") return # Use ExitStack to safely manage multiple output files with ExitStack() as stack: # Open all output files in write mode (auto-closes later) out_files = [stack.enter_context(open(fname, 'w')) for fname in args.outfile] # Process input file line by line (memory-friendly for large files) with open(args.infile, 'r') as infile: for line in infile: # Split line into columns (adjust delimiter if needed) columns = line.strip().split() # Write each selected column to its corresponding file for col_idx, out_file in zip(args.column, out_files): # Handle lines with fewer columns than requested if col_idx < len(columns): out_file.write(f"{columns[col_idx]}\n") else: # Write empty line if column doesn't exist in this row out_file.write("\n") if __name__ == '__main__': main()
Key Details & Usage:
- Argument Handling: The script uses Python's built-in
argparseto parse your command-line inputs. Thenargs='+'flag lets you pass multiple values for--columnand--outfile. - Validation: It first checks that the number of columns matches the number of output files—no more mismatched inputs!
- Safe File Management:
ExitStackensures all output files are closed properly, even if an error mid-execution. - Line-by-Line Processing: Reads the input file one line at a time, which is efficient for large files that might crash your RAM if loaded entirely.
Run it exactly like your example:
python septc.py --infile infile.txt --column 0 2 3 --outfile out1.txt out2.txt out3.txt
Customization Tips:
- Delimiter: If your file uses commas, tabs, or another separator instead of whitespace, replace
split()withsplit(',')(or your delimiter of choice). - Edge Case Adjustment: If you don't want empty lines for missing columns, you can replace the
elseblock withpassto skip writing anything for those lines.
内容的提问来源于stack exchange,提问作者nic
相关产品推荐
相关产品推荐

