Python新手求TXT转CSV实现指引:含自定义数据处理思路
Hey there! As someone who fumbled through similar file conversion tasks when I was starting out with Python, let me break down the key technical directions to help you get this done right.
Core Technical Steps to Follow
Here’s the breakdown of what you need to implement:
Read TXT Data in Blocks (Split by
#)
Your goal is to group lines until you hit a#—that group becomes one row of CSV data. You can do this by:- Opening the TXT file and reading line by line
- Accumulating lines into a temporary list until you encounter a
# - Processing that accumulated list as a single data entry, then resetting the list for the next block
- Don’t forget to handle the last block if your file doesn’t end with a
#!
Extract Data Using
=as Delimiter
For each line in your data block (likename=Aliceorage=30), split the line at the first=to separate the key (e.g.,name) and value (e.g.,Alice). Usingsplit("=", 1)ensures you don’t accidentally split values that contain=(likedescription=Python is fun=and useful).Map Data to CSV Columns
Define your target CSV columns upfront (e.g.,["name", "age", "city"]). Use a dictionary to map each extracted key to the corresponding CSV column—this way, even if the lines in your TXT block are out of order, the data ends up in the right CSV column.Write to CSV File
Use Python’s built-incsvmodule instead of manually writing comma-separated lines—it handles edge cases like values with commas or quotes automatically. You can usecsv.writerfor list-based rows, orcsv.DictWriterif you prefer working with dictionaries.
Example Implementation
Here’s a practical code snippet that ties all these steps together:
import csv # Define your target CSV columns (match the keys in your TXT data) csv_fields = ["name", "age", "city", "email"] # Initialize CSV file with headers (run this once at the start) with open("output.csv", "w", encoding="utf-8", newline="") as csv_file: writer = csv.writer(csv_file) writer.writerow(csv_fields) # Process the TXT file with open("input.txt", "r", encoding="utf-8") as txt_file: current_data_block = [] for line in txt_file: cleaned_line = line.strip() # Skip empty lines to avoid noise if not cleaned_line: continue if cleaned_line == "#": # Process the accumulated block if current_data_block: # Create a dictionary to hold data for this CSV row row_dict = {field: "" for field in csv_fields} for item in current_data_block: if "=" in item: # Split only at the first '=' to preserve values with '=' key, value = item.split("=", 1) key = key.strip() value = value.strip() # Only update if the key matches a CSV column if key in row_dict: row_dict[key] = value # Convert the dictionary to a list in CSV column order csv_row = [row_dict[field] for field in csv_fields] # Append to CSV with open("output.csv", "a", encoding="utf-8", newline="") as csv_file: writer = csv.writer(csv_file) writer.writerow(csv_row) # Reset for next block current_data_block = [] else: current_data_block.append(cleaned_line) # Handle the last block if the file doesn't end with '#' if current_data_block: row_dict = {field: "" for field in csv_fields} for item in current_data_block: if "=" in item: key, value = item.split("=", 1) key = key.strip() value = value.strip() if key in row_dict: row_dict[key] = value csv_row = [row_dict[field] for field in csv_fields] with open("output.csv", "a", encoding="utf-8", newline="") as csv_file: writer = csv.writer(csv_file) writer.writerow(csv_row)
Key Tips for Success
- Test with a small sample: Before processing your full TXT file, create a tiny test file with 2-3 blocks to verify the code works as expected.
- Handle encoding: Always specify
encoding="utf-8"when opening files to avoid weird character issues. - Debug with prints: Add
print(current_data_block)orprint(row_dict)to check if your data is being captured and mapped correctly. - Use
DictWriterfor simplicity: If you prefer working with dictionaries, replacecsv.writerwithcsv.DictWriter—it lets you write rows directly from therow_dictwithout converting to a list.
内容的提问来源于stack exchange,提问作者CodeModeOn

