如何在Python中匹配文本文件与Debian安全追踪JSON的包名?
问题描述
我是JSON新手,有一个每行存储一个包名的文本文件,想检查这些包名是否存在于Debian安全追踪的JSON数据中,并打印匹配的包名。我试了下面的代码,但没有任何输出:
def json_find(): json_file = json.dumps(info) with open("package_names.txt", "r") as f: for line in f: if line in json_file: print (line) json_find()
其中info已经加载了Debian安全追踪的JSON数据,但我没法正确遍历文本文件并在JSON里搜索包名。
文本文件示例:
nftables python3-translationstring gcc-8-base libpocojson60 passwd automake
JSON数据结构示例(包名为顶级键):
{ "389-ds-base": { "CVE-2012-0833": { "description": "The acllas__handle_group_entry function in servers/plugins/acl/acllas.c in 389 Directory Server before 1.2.10 does not properly handled access control instructions (ACIs) that use certificate groups, which allows remote authenticated LDAP users with a certificate group to cause a denial of service (infinite loop and CPU consumption) by binding to the server.", "scope": "local", "releases": { "bookworm": { "status": "resolved", "repositories": { "bookworm": "2.0.15-1" }, "fixed_version": "0", "urgency": "unimportant" } } }, "CVE-2012-2678": { "description": "389 Directory Server before 1.2.11.6 (aka Red Hat Directory Server before 8.2.10-3), after the password for a LDAP user has been changed and before the server has been reset, allows remote attackers to read the plaintext password via the unhashed#user#password attribute.", "scope": "local", "releases": { "bookworm": { "status": "resolved", "repositories": { "bookworm": "2.0.15-1" }, "fixed_version": "0", "urgency": "unimportant" } } } } }
比如,如果我的列表里有389-ds-base,希望能打印出这个包名。
问题原因与修复方案
原代码的问题
- 转字符串搜索既低效又易出错:
json.dumps(info)把JSON字典转成字符串后,用line in json_file匹配会因为文本行末尾的换行符(\n)导致匹配失败——比如文本里的nftables实际是nftables\n,而JSON字符串里的键是"nftables",完全不匹配。同时这种方式遍历整个字符串,数据量大时速度极慢。 - 没利用JSON的结构化特性:加载后的
info是Python字典,包名就是字典的顶级键,直接检查键是否存在才是正确做法。
修复后的代码
import json def json_find(): # 假设info已经通过json.load()或response.json()加载完成 # 把包名转成集合,提升查找效率 debian_packages = set(info.keys()) with open("package_names.txt", "r") as f: for line in f: # 去除每行的换行符和首尾空白 package = line.strip() # 直接检查包名是否在Debian包集合中 if package in debian_packages: print(package) json_find()
额外提示
- 如果是从网络获取JSON数据,用
requests.get(url).json()直接得到字典,无需手动处理字符串。 - 用集合存储包名是因为集合的成员检查比字典略快,数据量越大优势越明显。
内容的提问来源于stack exchange,提问作者Kobotecnico
相关产品推荐
相关产品推荐

