如何按起止模式合并文本文件多行并将短信数据清洗转换为CSV?
Got it, let's work through this problem together. The core challenge here is handling those multiline SMS messages—where a single text spans multiple lines in your .txt file—before splitting everything into the Date/Time/Name/Text columns for your CSV.
Step 1: Understand Your SMS File Structure
First, we need to identify what marks the start of a new SMS. For example, your .txt might look like this:
2024-05-20 14:30:00 [Alice] Hey there!
Did you see that new coffee shop downtown?
The lavender latte is amazing
2024-05-20 14:32:15 [Bob] Yeah, I went yesterday!
They also have homemade croissants that are to die for
In this case, every new message starts with a YYYY-MM-DD HH:MM:SS [Name] pattern. We'll use this pattern to detect when a new message begins, then merge all subsequent lines into a single text until the next message starts.
Step 2: Python Solution to Merge & Convert
Here's a practical script that handles merging multiline messages and writing the structured data to CSV. It uses regex to detect message starts and cleans up the content properly:
import csv import re # Regex pattern to match the start of a new SMS (adjust this if your format differs!) # This expects lines starting with "YYYY-MM-DD HH:MM:SS [Name]" message_start = re.compile(r'^\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2} \[.*?\]') merged_messages = [] current_msg = None # Read and preprocess the SMS file (clean empty lines and extra whitespace) with open('your_sms_file.txt', 'r', encoding='utf-8') as txt_file: cleaned_lines = [line.strip() for line in txt_file if line.strip()] for line in cleaned_lines: # Check if this line is the start of a new message if message_start.match(line): # Save the previous completed message if it exists if current_msg is not None: merged_messages.append(current_msg) # Split the start line into date, time, name, and initial text timestamp_part, rest = line.split(' [', 1) name, text = rest.split('] ', 1) current_msg = { 'date': timestamp_part.split(' ')[0], 'time': timestamp_part.split(' ')[1], 'name': name, 'text': text } else: # Append this line to the current message's text if current_msg is not None: current_msg['text'] += ' ' + line # Use space to join line breaks # Don't forget to add the last message to the list if current_msg is not None: merged_messages.append(current_msg) # Write the structured data to CSV with open('cleaned_sms.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=['Date', 'Time', 'Name', 'Text']) writer.writeheader() for msg in merged_messages: writer.writerow(msg)
Key Notes & Adjustments
- Adjust the regex pattern: If your SMS starts with a different format (e.g.,
2024/05/20 14:30 Alice:), tweak themessage_startregex to match your actual file structure. For example, user'^\d{4}/\d{2}/\d{2} \d{2}:\d{2} .+:'and adjust the splitting logic accordingly. - Encoding: Using
utf-8ensures special characters (emojis, accented letters) are preserved correctly. - Test with a small sample: Run the script on a tiny portion of your .txt first to verify merging works as expected—this helps catch edge cases like messages with unusual line breaks.
内容的提问来源于stack exchange,提问作者Eoin

