You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将含缺失字段的聊天记录TXT文件转换为Pandas DataFrame?

Alright, let's break down how to turn this tricky TXT file into a clean pandas DataFrame. I’ve dealt with similar messy structured files before, so here’s a step-by-step solution that handles both the named fields and the unlabeled chat history:

Step 1: Read and Split the File into Individual Records

First, we’ll read the entire file and split it into separate chat records using that long string of equals signs as the delimiter. We’ll also filter out any empty blocks that might show up at the start/end of the file.

Step 2: Parse Each Record into Fields + Chat History

For each record block, we’ll loop through every line:

  • Lines that match your predefined field names (like Timestamp or Visitor ID) get split into key-value pairs.
  • Any lines that don’t match a known field get grouped together as the chat history for that record.

Full Code Implementation

import pandas as pd

# Define the separator that splits individual chat records
RECORD_SEPARATOR = "================================================================================"

# Read the entire file content
with open("your_chat_records.txt", "r", encoding="utf-8") as file:
    raw_content = file.read()

# Split into records and remove empty/whitespace-only blocks
chat_records = [block.strip() for block in raw_content.split(RECORD_SEPARATOR) if block.strip()]

# Prepare a list to store parsed data
parsed_data = []

# List of all your known field names (exact matches required!)
known_fields = {
    "Timestamp", "Unread", "Visitor ID", "Visitor Name", "Visitor Email",
    "Visitor Notes", "I\\P", "Country Code", "Country Name", "Region",
    "City", "User Agent", "Platform", "Browser"
}

for record in chat_records:
    # Split the record into lines, skipping empty ones
    lines = [line.strip() for line in record.split("\n") if line.strip()]
    
    record_dict = {}
    chat_history = []
    
    for line in lines:
        # Check if the line starts with a known field (adjust split logic if your file uses colons!)
        # Example: If your line is "Timestamp: 2021-06-09...", use split(": ", maxsplit=1) instead
        field_matched = False
        for field in known_fields:
            if line.startswith(field):
                # Split field name from value (only split once to preserve spaces in values)
                _, value = line.split(maxsplit=1)
                record_dict[field] = value.strip()
                field_matched = True
                break
        
        if not field_matched:
            # Add non-field lines to chat history
            chat_history.append(line)
    
    # Add the combined chat history to the record
    record_dict["Chat History"] = "\n".join(chat_history)
    parsed_data.append(record_dict)

# Convert the list of dictionaries to a DataFrame
chat_df = pd.DataFrame(parsed_data)

# Optional: Reorder columns to match your preferred sequence
desired_column_order = [
    "Timestamp", "Unread", "Visitor ID", "Visitor Name", "Visitor Email",
    "Visitor Notes", "I\\P", "Country Code", "Country Name", "Region",
    "City", "User Agent", "Platform", "Browser", "Chat History"
]
chat_df = chat_df[desired_column_order]

# Preview the result
print(chat_df.head())

Key Adjustments for Your File

  • Field-Value Separator: If your TXT uses colons (e.g., Timestamp: 2021-06-09...) instead of just spaces, replace the split(maxsplit=1) line with field, value = line.split(": ", maxsplit=1) and remove the loop over known_fields—just check if the extracted field is in known_fields.
  • Encoding: If your file has special characters, tweak the encoding parameter in open() (e.g., encoding="latin-1").
  • Field Name Exactness: Make sure the known_fields set matches the exact spelling/capitalization in your TXT file (including that I\P field—we escaped the backslash with \\ to handle it correctly in Python).

Once you run this, you’ll have a DataFrame where each row represents a full chat record, with all your named fields as columns and the unlabeled chat history neatly stored in its own Chat History column.

内容的提问来源于stack exchange,提问作者Karim Jr.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 19:43:09