You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3解析日志中JSON字符串失败问题求助

问题排查与解决方案

问题背景

从多个日志文件收集JSON字符串存入列表,遍历发现这些JSON始终是字符串类型,无法转为Python可操作的列表/字典。这些字符串用jq能正常解析,在ipython3赋值变量也能得到预期列表,但Python代码里用json.loads()报JSONDecodeError,试json.load()又触发AttributeError(提示'str'对象无'read'属性),尝试去除字符串首尾括号也没解决。

JSON示例

{
  "system_name": "orderxxx-live",
  "manifest_assets": [
    "xx-xxx",
    "xx-xx",
    "xx-xxx",
    "xx-xxx",
    "xx-xxx",
    "xx-xxx",
    "xx-xxx",
    "xx-xxx",
    "xx-xxx"
  ],
  "certificate_expiration_date": "2025-08-23 08:07:33",
  "artifacts": [
    {
      "artifact_name": "item-6.0.war",
      "artifact_version": "6.0.438"
    },
    {
      "artifact_name": "service-v1.0.war",
      "artifact_version": "2.0.10"
    },
    {
      "artifact_name": "camunda-webapp-ee-jboss-7.16.7-ee.war",
      "artifact_version": "7.16.7-ee"
    },
    {
      "artifact_name": "callback-v7.0.ear",
      "artifact_version": "1.0.56"
    }
  ],
  "application_server": [
    "postgres_dbms: localhost:5432/orderxxx?currentSchema=camunda",
    "applications_server: WildFly 34.0",
    "java_path: /usr/lib/jvm/temurin-11-jdk-amd64/bin/java"
  ],
  "additional_data": []
}

初始Python代码

def __main__():
  sort_list = []

  ## extract data from logfiles
  for file_name in glob.glob(logstore_path + '/*/*/system/the_log.log'):
    with open(file_name, 'r') as ASS:
      for line in ASS:
        sort_list.append(line)

  sort_list = list(set(sort_list))
  for i in sort_list:
    print(type(i)) ## here only strings
    print(json.loads(i))

初始报错信息

Traceback (most recent call last):
  File "/home/druuhl/git/omdm_app_montool_apache/montool_vhost1/t_i_g/./tig_glue_data.py", line 41, in <module>
    __main__()
    ~~~~~~~~^^
  File "/home/druuhl/git/omdm_app_montool_apache/montool_vhost1/t_i_g/./tig_glue_data.py", line 30, in __main__
    ttt = json.loads(ttt)
  File "/usr/lib/python3.13/json/__init__.py", line 346, in loads
    return _default_decoder.decode(s)
           ~~~~~~~~~~~~~~~~~~~~~~~^^^
  File "/usr/lib/python3.13/json/decoder.py", line 345, in decode
    obj, end = self.raw_decode(s, idx=_w(s, 0).end())
               ~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/lib/python3.13/json/decoder.py", line 363, in raw_decode
    raise JSONDecodeError("Expecting value", s, err.value) from None
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)

调整后的代码

def __main__():
  sort_list = []

  ## extract data from logfiles
  for file_name in glob.glob(logstore_path + '/*/*/system/t_i_g.log'):
    with open(file_name, 'r') as ASS:
      for line in ASS:
        sort_list.append(line)

  sort_list = list(set(sort_list))

  #print(sort_list)
  for i in sort_list:
    print('foo')
    print(i)
    ## one output with and one without brackets
    i = i[1:-1]
    i = i[:-1]
    print(i)
    print(json.load(i))

调整后的报错信息

./tig_glue_data.py
foo
[{"system_name": "migrationentry-dev", "manifest_assets": [], "certificate_expiration_date": "2026-01-18 10:59:47", "artifacts": [{"artifact_name": ..... "}]

Traceback (most recent call last):
  File "/home/druuhl/git/omdm_app_montool_apache/montool_vhost1/t_i_g/./tig_glue_data.py", line 45, in <module>
    __main__()
    ~~~~~~~~^^
  File "/home/druuhl/git/omdm_app_montool_apache/montool_vhost1/t_i_g/./tig_glue_data.py", line 29, in __main__
    print(json.load(i))
          ~~~~~~~~~^^^
  File "/usr/lib/python3.13/json/__init__.py", line 293, in load
    return loads(fp.read(),
                 ^^^^^^^
AttributeError: 'str' object has no attribute 'read'

问题原因分析

  1. 初始JSONDecodeError原因:
    • 读取日志行时,每行末尾带有换行符\n,或者存在空行、非JSON格式的行,导致json.loads()解析失败。
    • 使用list(set(sort_list))去重时,可能引入了无效的空字符串或格式损坏的行。
  2. 调整后AttributeError原因:
    • json.load()接收的是文件对象,而非字符串,直接传入字符串会触发该错误,应该用json.loads()处理字符串。

解决方案

方案1:清理字符串并过滤无效行

import json
import glob

def __main__():
    sort_list = []
    logstore_path = "/path/to/your/logs"  # 替换为实际路径

    # 读取日志文件,过滤无效行并清理字符串
    for file_name in glob.glob(logstore_path + '/*/*/system/the_log.log'):
        with open(file_name, 'r') as ASS:
            for line in ASS:
                stripped_line = line.strip()  # 去除首尾空白字符(包括换行符)
                if stripped_line:  # 跳过空行
                    sort_list.append(stripped_line)

    # 去重并解析
    unique_lines = list(set(sort_list))
    for line in unique_lines:
        try:
            data = json.loads(line)
            print(type(data))  # 应为dict或list
            print(data)
        except json.JSONDecodeError as e:
            print(f"解析失败的行: {line}")
            print(f"错误信息: {e}")

if __name__ == "__main__":
    __main__()

方案2:处理包裹在数组中的JSON(根据调整后代码的输出)

如果日志中的JSON是被数组包裹的(比如[{"system_name": ...}]),解析时无需手动去除括号,直接用json.loads()即可:

# 在解析部分替换为:
try:
    data = json.loads(line)
    # 如果是数组,遍历内部元素
    if isinstance(data, list):
        for item in data:
            print(item)
    else:
        print(data)
except json.JSONDecodeError as e:
    print(f"解析失败: {e}")

关键注意点

  • 区分json.load()和json.loads():load()用于读取文件对象,loads()用于解析字符串。
  • 清理字符串:读取行时必须用strip()去除换行符和首尾空白,否则会导致解析失败。
  • 过滤无效行:跳过空行或非JSON格式的行,避免解析报错。

内容的提问来源于stack exchange,提问作者druuhl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 20:23:14