如何将特定格式文本文件解析并转换为CSV格式?
It sounds like your current approach of replacing colons globally is causing more harm than good—this breaks the critical field-value separation you need. Let’s fix this with a targeted approach using Python’s re module to extract exact data pairs, then leverage the built-in csv module to generate a properly structured CSV.
Step 1: Clarify the Input Structure
First, let’s assume your filtered text (after stripping email headers) looks something like this (adjust to match your actual data):
TICKET NUMBER: 12345
OLD SOURCE IPIP: 192.168.1.100
NEW SOURCE IP: 10.0.0.50TICKET NUMBER: 12346
OLD SOURCE IPIP: 172.16.3.20
STATUS: Resolved
Each ticket record is separated by a divider (like --- or blank lines), and each line follows a Field Name: Value format.
Step 2: Extract Records and Key-Value Pairs
Instead of modifying colons directly, we’ll use regex to capture field names and their values, then group these into individual ticket records. Here’s a sample code snippet:
import re import csv # Read your filtered text file with open('filtered_tickets.txt', 'r') as f: text = f.read() # Split text into individual ticket records (adjust separator regex as needed) # Use r'\n\s*\n' instead of r'---' if records are separated by blank lines records = re.split(r'---', text) # Regex pattern to match field-value pairs: captures multi-word fields and their values field_pattern = re.compile(r'^(\w+(?:\s+\w+)*):\s*(.*)$', re.MULTILINE) # Collect all unique fields for CSV headers and store ticket data all_fields = set() ticket_data = [] for record in records: # Skip empty records from leading/trailing separators if not record.strip(): continue # Extract all key-value pairs from the current ticket pairs = field_pattern.findall(record) ticket_dict = {key.strip(): value.strip() for key, value in pairs} ticket_data.append(ticket_dict) # Track all unique fields to build consistent headers all_fields.update(ticket_dict.keys()) # Convert fields to a sorted list for consistent CSV columns headers = sorted(all_fields) # Write to properly formatted CSV with open('tickets.csv', 'w', newline='') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=headers) writer.writeheader() writer.writerows(ticket_data)
Step 3: Why This Solves Your Problem
- Precise Regex Capture: The pattern
^(\w+(?:\s+\w+)*):\s*(.*)$correctly identifies multi-word fields likeTICKET NUMBERand separates them from their values, avoiding the colon-replacement mess from your original code. - Record Grouping: Splitting the text into individual tickets ensures each row in your CSV corresponds to one ticket.
- Robust CSV Handling: Using
csv.DictWriterautomatically maps each ticket’s data to the correct columns, even if some tickets are missing certain fields (those cells will be left empty instead of breaking the structure).
Customization Tips
- If your field names include special characters (like hyphens), update the regex to
^([\w\s-]+):\s*(.*)$. - If records aren’t separated by
---, adjust there.splitpattern (e.g.,r'\n{2,}'for two or more blank lines). - If values span multiple lines, modify the regex to capture multi-line content (use a non-greedy match and look ahead for the next field or record separator).
内容的提问来源于stack exchange,提问作者Happy.Hartman

