使用Pandas处理Symphony日志时解析畸形JSON遇到问题
Got it, let's walk through how to fix this messy Symphony log JSON so Pandas can handle it. First, looking at your sample line, there are a few obvious issues breaking valid JSON structure:
- None of the keys (like
id,work_assignment,declaration) are wrapped in double quotes—JSON requires keys to be quoted strings. - The
work_assignmentvalue uses=>which isn't valid JSON syntax (I’m guessing this is a custom separator for a transition, like 43313 to 43313). - The
declarationfield has unquoted HTML content with line breaks, plus a random trailing=>that cuts off the line prematurely.
Here's a practical, step-by-step solution to clean these logs:
Step 1: Write a Custom Cleaning Function
We’ll use regex and string manipulation to fix each of these issues. Let’s build a function tailored to your log format:
import re import json import pandas as pd def clean_symphony_log(line): # Wrap all unquoted keys in double quotes line = re.sub(r'(?<!")(\b\w+\b)(?=:)', r'"\1"', line) # Handle work_assignment's => by quoting the entire value (since it's a transition) line = re.sub(r'"work_assignment":(\d+=>\d+)', r'"work_assignment":"\1"', line) # Quote the declaration HTML content, stripping any trailing invalid => def quote_html_declaration(match): key = match.group(1) # Capture everything after declaration: until the next quoted key or end of line value = match.group(2).rstrip('=>').strip() # Escape double quotes inside HTML to avoid breaking JSON escaped_value = value.replace('"', '\\"') return f'"{key}": "{escaped_value}"' # Use DOTALL to match across line breaks in the HTML line = re.sub(r'(\bdeclaration\b):(.*?)(?=\s*"\w+"|:|$)', quote_html_declaration, line, flags=re.DOTALL) # Fix truncated lines by adding a closing brace if missing if not line.strip().endswith('}'): line = line.rstrip('=>').strip() + '}' return line
Step 2: Test the Cleaner on Your Sample Line
Let’s run this function on your example log to see the result:
sample_log = '{id:46025, work_assignment:43313=>43313, declaration:<p><strong>Bijkomende interventie.</strong></p>\r\n\r\n<p>H </p>\r\n\r\n<p><strong><em>Vaststellingen.</em></strong></p>\r\n\r\n<p><strong><em>CV. </em></strong>De.</p>=><p>...' cleaned_json = clean_symphony_log(sample_log) print(cleaned_json)
This should output a valid JSON string like:
{"id":46025, "work_assignment":"43313=>43313", "declaration": "<p><strong>Bijkomende interventie.</strong></p>\r\n\r\n<p>H </p>\r\n\r\n<p><strong><em>Vaststellingen.</em></strong></p>\r\n\r\n<p><strong><em>CV. </em></strong>De.</p>"}
Step 3: Parse Cleaned Logs into a Pandas DataFrame
Once you’ve cleaned all your log lines, load them into a DataFrame. If your logs are in a text file, here’s how to process them:
# Load and process each line from the log file cleaned_logs = [] with open('symphony_logs.txt', 'r', encoding='utf-8') as log_file: for line in log_file: cleaned_line = clean_symphony_log(line.strip()) try: # Parse cleaned JSON into a dictionary log_dict = json.loads(cleaned_line) cleaned_logs.append(log_dict) except json.JSONDecodeError as e: # Catch lines that still fail to parse (adjust the cleaner if needed) print(f"Skipping invalid line: {line[:50]}... Error: {e}") continue # Convert to DataFrame df = pd.DataFrame(cleaned_logs) print(df.head())
Quick Adjustments to Note
- If
=>appears elsewhere in your logs with different meanings, tweak the regex to target only thework_assignmentfield. - If your HTML has other tricky characters (like single quotes), add extra escaping steps as needed.
- For severely truncated lines, you may need to add additional checks to handle partial JSON structures.
Content of the question originates from Stack Exchange, question author Maarten Kesselaers

