如何用Python批量给文本文件添加表头并整合至CSV数据库?
Python Solution for Batch Processing TXT Files into Structured CSV
Absolutely! Python is an excellent tool for this task—we can build a script that iterates through your TXT files, prompts you to input the required metadata for each one, and compiles everything into the CSV format you need.
Here's a step-by-step implementation:
Approach Overview
- Locate TXT Files: Use the
osmodule to find all.txtfiles in a specified directory. - Collect Metadata: For each file, prompt you to enter the
doc_id,speech_type,author, anddatefields manually. - Read File Content: Extract the full text from each TXT file to use as the
textcolumn value. - Write to CSV: Use Python's built-in
csvmodule to append each entry to the output CSV, following your desired format.
Complete Code Example
import os import csv def process_txt_to_csv(input_dir, output_csv): # Define the CSV header matching your desired format csv_headers = ["doc_id", "speech_type", "author", "date", "text"] # Open the output CSV file in write mode (creates a new file or overwrites existing) with open(output_csv, mode='w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter( csv_file, fieldnames=csv_headers, delimiter=';', quotechar='"', quoting=csv.QUOTE_ALL ) writer.writeheader() # Iterate over all files in the input directory for filename in os.listdir(input_dir): if filename.endswith('.txt'): file_path = os.path.join(input_dir, filename) print(f"\nProcessing file: {filename}") # Prompt user for metadata (adjust prompts to match your exact needs) doc_id = input("Enter doc_id: ").strip() speech_type = input("Enter speech_type (e.g., speech): ").strip() author = input("Enter author (e.g., speaker): ").strip() date = input("Enter date (yyyy-mm-dd): ").strip() # Read the content of the TXT file try: with open(file_path, 'r', encoding='utf-8') as txt_file: text = txt_file.read().strip() except Exception as e: print(f"Error reading {filename}: {str(e)}") continue # Write the entry to CSV writer.writerow({ "doc_id": doc_id, "speech_type": speech_type, "author": author, "date": date, "text": text }) print(f"Successfully added entry for {filename}") # Configuration - adjust these paths to match your setup INPUT_DIRECTORY = "./your_txt_files" # Replace with your folder containing TXT files OUTPUT_CSV = "./speeches.csv" # Replace with your desired output CSV path # Run the process if __name__ == "__main__": process_txt_to_csv(INPUT_DIRECTORY, OUTPUT_CSV) print("\nAll files processed! Output saved to:", OUTPUT_CSV)
Key Notes & Enhancements
- Encoding: The script uses
utf-8encoding for both reading TXT files and writing CSV. If your files use a different encoding (likegbk), adjust theencodingparameter in theopen()calls to match. - Metadata Persistence: If you don't want to re-enter metadata for files you've already processed, you could add a check for existing entries in the CSV or save metadata to a separate JSON file for reuse.
- Batch Metadata Input: If you have metadata stored in a separate file (e.g., a CSV mapping filenames to their metadata), you can modify the script to read from that instead of prompting manually—just let me know if you need help with that!
- Error Handling: The script includes basic error handling for file reading, but you can expand it to handle edge cases like empty files or invalid input formats.
To use this script:
- Replace
INPUT_DIRECTORYwith the path to your folder containing all TXT files. - Set
OUTPUT_CSVto your desired output file path. - Run the script—you'll be prompted to enter metadata for each TXT file one by one.
内容的提问来源于stack exchange,提问作者user3081750
相关产品推荐
相关产品推荐

