如何将含缺失字段的聊天记录TXT文件转换为Pandas DataFrame?
Alright, let's break down how to turn this tricky TXT file into a clean pandas DataFrame. I’ve dealt with similar messy structured files before, so here’s a step-by-step solution that handles both the named fields and the unlabeled chat history:
Step 1: Read and Split the File into Individual Records
First, we’ll read the entire file and split it into separate chat records using that long string of equals signs as the delimiter. We’ll also filter out any empty blocks that might show up at the start/end of the file.
Step 2: Parse Each Record into Fields + Chat History
For each record block, we’ll loop through every line:
- Lines that match your predefined field names (like
TimestamporVisitor ID) get split into key-value pairs. - Any lines that don’t match a known field get grouped together as the chat history for that record.
Full Code Implementation
import pandas as pd # Define the separator that splits individual chat records RECORD_SEPARATOR = "================================================================================" # Read the entire file content with open("your_chat_records.txt", "r", encoding="utf-8") as file: raw_content = file.read() # Split into records and remove empty/whitespace-only blocks chat_records = [block.strip() for block in raw_content.split(RECORD_SEPARATOR) if block.strip()] # Prepare a list to store parsed data parsed_data = [] # List of all your known field names (exact matches required!) known_fields = { "Timestamp", "Unread", "Visitor ID", "Visitor Name", "Visitor Email", "Visitor Notes", "I\\P", "Country Code", "Country Name", "Region", "City", "User Agent", "Platform", "Browser" } for record in chat_records: # Split the record into lines, skipping empty ones lines = [line.strip() for line in record.split("\n") if line.strip()] record_dict = {} chat_history = [] for line in lines: # Check if the line starts with a known field (adjust split logic if your file uses colons!) # Example: If your line is "Timestamp: 2021-06-09...", use split(": ", maxsplit=1) instead field_matched = False for field in known_fields: if line.startswith(field): # Split field name from value (only split once to preserve spaces in values) _, value = line.split(maxsplit=1) record_dict[field] = value.strip() field_matched = True break if not field_matched: # Add non-field lines to chat history chat_history.append(line) # Add the combined chat history to the record record_dict["Chat History"] = "\n".join(chat_history) parsed_data.append(record_dict) # Convert the list of dictionaries to a DataFrame chat_df = pd.DataFrame(parsed_data) # Optional: Reorder columns to match your preferred sequence desired_column_order = [ "Timestamp", "Unread", "Visitor ID", "Visitor Name", "Visitor Email", "Visitor Notes", "I\\P", "Country Code", "Country Name", "Region", "City", "User Agent", "Platform", "Browser", "Chat History" ] chat_df = chat_df[desired_column_order] # Preview the result print(chat_df.head())
Key Adjustments for Your File
- Field-Value Separator: If your TXT uses colons (e.g.,
Timestamp: 2021-06-09...) instead of just spaces, replace thesplit(maxsplit=1)line withfield, value = line.split(": ", maxsplit=1)and remove the loop overknown_fields—just check if the extractedfieldis inknown_fields. - Encoding: If your file has special characters, tweak the
encodingparameter inopen()(e.g.,encoding="latin-1"). - Field Name Exactness: Make sure the
known_fieldsset matches the exact spelling/capitalization in your TXT file (including thatI\Pfield—we escaped the backslash with\\to handle it correctly in Python).
Once you run this, you’ll have a DataFrame where each row represents a full chat record, with all your named fields as columns and the unlabeled chat history neatly stored in its own Chat History column.
内容的提问来源于stack exchange,提问作者Karim Jr.

