Python文本转JSON代码输出异常,求正确实现方案
文本转JSON格式的代码修正问题
问题描述
需要将指定格式的文本转换为JSON数组格式,现有代码输出不符合预期,请求修正。
待转换文本
Reason: RSA private key Date: 2021-05-10 01:37:04 Hash: 55cd03eb0f76c7b23d6c53dd2555f83709f834 Filepath: labs/lab-files/open-ssl-lab-files/private-key.pem Branch: origin/master Commit: 2.8.45: Updated Docker Files for PHP 8 -----BEGIN RSA PRIVATE KEY----- Reason: Password in URL Date: 2018-09-28 08:18:39 Hash: 0f9d4a66f2540f5154c01476e811615aad6291 Filepath: includes/hints/sqlmap-hint.inc Branch: origin/master Commit: =2.6.67 mysql://root:""@192.168.56.102:5123/nowasp Reason: Password in URL Date: 2018-09-28 08:18:39 Hash: 0f9d4a66f2540f5154c01476e811615aad6291 Filepath: owasp-esapi-php/lib/simpletest/docs/en/authentication_documentation.html Branch: origin/master Commit: =2.6.67 http://<strong>Me:Secret@</strong>www.lastcraft.com/protected/')
期望输出JSON数组
[ { "Reason": "RSA private key", "Date": "2021-05-10 01:37:04", "Hash": "55cd03eb0f76c7b23d6c53dd2555f83709f834", "Filepath": "labs/lab-files/open-ssl-lab-files/private-key.pem", "Branch": "origin/master", "Commit": "2.8.45: Updated Docker Files for PHP 8", "url":"-----BEGIN RSA PRIVATE KEY-----"}, { "Reason": "Password in URL", "Date": "2018-09-28 08:18:39", "Hash": "0f9d4a66f2540f5154c01476e811615aad6291", "Filepath": "includes/hints/sqlmap-hint.inc", "Branch": "origin/master", "Commit": "=2.6.67", "url":"mysql://root:""@192.168.56.102:5123/nowasp" }, { "Reason": "Password in URL", "Date": "2018-09-28 08:18:39", "Hash": "0f9d4a66f2540f5154c01476e811615aad6291", "Filepath":"owasp-esapi-php/lib/simpletest/docs/en/authentication_documentation.html", "Branch": "origin/master", "Commit": "=2.6.67", "url":"http://<strong>Me:Secret@</strong>www.lastcraft.com/protected/');'" } ]
现有错误代码
import re h = [] dy = {} for index,name in enumerate(thy): if (bool(re.search('://',name))) == True: dy.update({'url':name}) continue elif (bool(re.search('-----BEGIN RSA',name))) == True: dy.update({'url':name}) continue elif index % 6 == 0: h.append(dy) print(dy) else: a = name.split(':',1)[0] b = name.split(':',1)[1] dy.update({a:b})
代码问题分析
- 条目分割逻辑错误:
index % 6 == 0的判断时机不对,无法正确识别每个条目的结束点。 - 未处理空行:原文本中的空行是条目分隔符,会导致
split操作报错。 - 字典未重置:所有条目共用同一个
dy字典,最终列表中所有元素都会被最后一个条目覆盖。 - 缺少条目完成后的收集逻辑:没有在每个条目结束时将字典加入列表并重置。
修正后的代码
import re import json # 预处理文本行:过滤空行,去除每行首尾空格 thy = [line.strip() for line in thy if line.strip()] result = [] current_entry = {} for line in thy: # 匹配URL或RSA私钥开头行,作为当前条目的url字段 if re.search(r'://', line) or re.search(r'-----BEGIN RSA', line): current_entry['url'] = line # 将当前条目加入结果列表,重置字典准备下一个条目 result.append(current_entry) current_entry = {} else: # 处理键值对行,分割键和值并去除首尾空格 if ':' in line: key, value = line.split(':', 1) current_entry[key.strip()] = value.strip() # 生成格式化的JSON输出 print(json.dumps(result, indent=2))
代码说明
- 文本预处理:先过滤空行、去除每行首尾空格,避免无效内容干扰解析。
- 条目识别:当遇到URL或RSA私钥行时,判定当前条目收集完成,将字典加入结果列表并重置。
- 键值对解析:对普通行按
:分割键和值,确保键值对格式正确存入字典。 - JSON格式化:使用
json.dumps生成带缩进的JSON字符串,完全匹配期望输出格式。
内容的提问来源于stack exchange,提问作者soumya ranjan
相关产品推荐
相关产品推荐

