You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas处理Symphony日志时解析畸形JSON遇到问题

Fixing Malformed Symphony Log JSON for Pandas Parsing

Got it, let's walk through how to fix this messy Symphony log JSON so Pandas can handle it. First, looking at your sample line, there are a few obvious issues breaking valid JSON structure:

  • None of the keys (like id, work_assignment, declaration) are wrapped in double quotes—JSON requires keys to be quoted strings.
  • The work_assignment value uses => which isn't valid JSON syntax (I’m guessing this is a custom separator for a transition, like 43313 to 43313).
  • The declaration field has unquoted HTML content with line breaks, plus a random trailing => that cuts off the line prematurely.

Here's a practical, step-by-step solution to clean these logs:

Step 1: Write a Custom Cleaning Function

We’ll use regex and string manipulation to fix each of these issues. Let’s build a function tailored to your log format:

import re
import json
import pandas as pd

def clean_symphony_log(line):
    # Wrap all unquoted keys in double quotes
    line = re.sub(r'(?<!")(\b\w+\b)(?=:)', r'"\1"', line)
    
    # Handle work_assignment's => by quoting the entire value (since it's a transition)
    line = re.sub(r'"work_assignment":(\d+=>\d+)', r'"work_assignment":"\1"', line)
    
    # Quote the declaration HTML content, stripping any trailing invalid =>
    def quote_html_declaration(match):
        key = match.group(1)
        # Capture everything after declaration: until the next quoted key or end of line
        value = match.group(2).rstrip('=>').strip()
        # Escape double quotes inside HTML to avoid breaking JSON
        escaped_value = value.replace('"', '\\"')
        return f'"{key}": "{escaped_value}"'
    
    # Use DOTALL to match across line breaks in the HTML
    line = re.sub(r'(\bdeclaration\b):(.*?)(?=\s*"\w+"|:|$)', quote_html_declaration, line, flags=re.DOTALL)
    
    # Fix truncated lines by adding a closing brace if missing
    if not line.strip().endswith('}'):
        line = line.rstrip('=>').strip() + '}'
    
    return line

Step 2: Test the Cleaner on Your Sample Line

Let’s run this function on your example log to see the result:

sample_log = '{id:46025, work_assignment:43313=>43313, declaration:<p><strong>Bijkomende interventie.</strong></p>\r\n\r\n<p>H&nbsp;</p>\r\n\r\n<p><strong><em>Vaststellingen.</em></strong></p>\r\n\r\n<p><strong><em>CV. </em></strong>De.</p>=><p>...'

cleaned_json = clean_symphony_log(sample_log)
print(cleaned_json)

This should output a valid JSON string like:

{"id":46025, "work_assignment":"43313=>43313", "declaration": "<p><strong>Bijkomende interventie.</strong></p>\r\n\r\n<p>H&nbsp;</p>\r\n\r\n<p><strong><em>Vaststellingen.</em></strong></p>\r\n\r\n<p><strong><em>CV. </em></strong>De.</p>"}

Step 3: Parse Cleaned Logs into a Pandas DataFrame

Once you’ve cleaned all your log lines, load them into a DataFrame. If your logs are in a text file, here’s how to process them:

# Load and process each line from the log file
cleaned_logs = []
with open('symphony_logs.txt', 'r', encoding='utf-8') as log_file:
    for line in log_file:
        cleaned_line = clean_symphony_log(line.strip())
        try:
            # Parse cleaned JSON into a dictionary
            log_dict = json.loads(cleaned_line)
            cleaned_logs.append(log_dict)
        except json.JSONDecodeError as e:
            # Catch lines that still fail to parse (adjust the cleaner if needed)
            print(f"Skipping invalid line: {line[:50]}... Error: {e}")
            continue

# Convert to DataFrame
df = pd.DataFrame(cleaned_logs)
print(df.head())

Quick Adjustments to Note

  • If => appears elsewhere in your logs with different meanings, tweak the regex to target only the work_assignment field.
  • If your HTML has other tricky characters (like single quotes), add extra escaping steps as needed.
  • For severely truncated lines, you may need to add additional checks to handle partial JSON structures.

Content of the question originates from Stack Exchange, question author Maarten Kesselaers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:41:31