You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按起止模式合并文本文件多行并将短信数据清洗转换为CSV?

How to Merge Multiline SMS Content and Convert to Structured CSV

Got it, let's work through this problem together. The core challenge here is handling those multiline SMS messages—where a single text spans multiple lines in your .txt file—before splitting everything into the Date/Time/Name/Text columns for your CSV.

Step 1: Understand Your SMS File Structure

First, we need to identify what marks the start of a new SMS. For example, your .txt might look like this:

2024-05-20 14:30:00 [Alice] Hey there!
Did you see that new coffee shop downtown?
The lavender latte is amazing
2024-05-20 14:32:15 [Bob] Yeah, I went yesterday!
They also have homemade croissants that are to die for

In this case, every new message starts with a YYYY-MM-DD HH:MM:SS [Name] pattern. We'll use this pattern to detect when a new message begins, then merge all subsequent lines into a single text until the next message starts.

Step 2: Python Solution to Merge & Convert

Here's a practical script that handles merging multiline messages and writing the structured data to CSV. It uses regex to detect message starts and cleans up the content properly:

import csv
import re

# Regex pattern to match the start of a new SMS (adjust this if your format differs!)
# This expects lines starting with "YYYY-MM-DD HH:MM:SS [Name]"
message_start = re.compile(r'^\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2} \[.*?\]')

merged_messages = []
current_msg = None

# Read and preprocess the SMS file (clean empty lines and extra whitespace)
with open('your_sms_file.txt', 'r', encoding='utf-8') as txt_file:
    cleaned_lines = [line.strip() for line in txt_file if line.strip()]

    for line in cleaned_lines:
        # Check if this line is the start of a new message
        if message_start.match(line):
            # Save the previous completed message if it exists
            if current_msg is not None:
                merged_messages.append(current_msg)
            
            # Split the start line into date, time, name, and initial text
            timestamp_part, rest = line.split(' [', 1)
            name, text = rest.split('] ', 1)
            current_msg = {
                'date': timestamp_part.split(' ')[0],
                'time': timestamp_part.split(' ')[1],
                'name': name,
                'text': text
            }
        else:
            # Append this line to the current message's text
            if current_msg is not None:
                current_msg['text'] += ' ' + line  # Use space to join line breaks

    # Don't forget to add the last message to the list
    if current_msg is not None:
        merged_messages.append(current_msg)

# Write the structured data to CSV
with open('cleaned_sms.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=['Date', 'Time', 'Name', 'Text'])
    writer.writeheader()
    for msg in merged_messages:
        writer.writerow(msg)

Key Notes & Adjustments

  • Adjust the regex pattern: If your SMS starts with a different format (e.g., 2024/05/20 14:30 Alice:), tweak the message_start regex to match your actual file structure. For example, use r'^\d{4}/\d{2}/\d{2} \d{2}:\d{2} .+:' and adjust the splitting logic accordingly.
  • Encoding: Using utf-8 ensures special characters (emojis, accented letters) are preserved correctly.
  • Test with a small sample: Run the script on a tiny portion of your .txt first to verify merging works as expected—this helps catch edge cases like messages with unusual line breaks.

内容的提问来源于stack exchange,提问作者Eoin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:13:06