You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中匹配文本文件与Debian安全追踪JSON的包名?

问题描述

我是JSON新手,有一个每行存储一个包名的文本文件,想检查这些包名是否存在于Debian安全追踪的JSON数据中,并打印匹配的包名。我试了下面的代码,但没有任何输出:

def json_find():
    json_file = json.dumps(info)
    with open("package_names.txt", "r") as f:
       for line in f:
            if line in json_file:
                print (line)
json_find()

其中info已经加载了Debian安全追踪的JSON数据,但我没法正确遍历文本文件并在JSON里搜索包名。

文本文件示例:

nftables
python3-translationstring
gcc-8-base
libpocojson60
passwd
automake

JSON数据结构示例(包名为顶级键):

{
  "389-ds-base": {
    "CVE-2012-0833": {
      "description": "The acllas__handle_group_entry function in servers/plugins/acl/acllas.c in 389 Directory Server before 1.2.10 does not properly handled access control instructions (ACIs) that use certificate groups, which allows remote authenticated LDAP users with a certificate group to cause a denial of service (infinite loop and CPU consumption) by binding to the server.",
      "scope": "local",
      "releases": {
        "bookworm": {
          "status": "resolved",
          "repositories": {
            "bookworm": "2.0.15-1"
          },
          "fixed_version": "0",
          "urgency": "unimportant"
        }
      }
    },
    "CVE-2012-2678": {
      "description": "389 Directory Server before 1.2.11.6 (aka Red Hat Directory Server before 8.2.10-3), after the password for a LDAP user has been changed and before the server has been reset, allows remote attackers to read the plaintext password via the unhashed#user#password attribute.",
      "scope": "local",
      "releases": {
        "bookworm": {
          "status": "resolved",
          "repositories": {
            "bookworm": "2.0.15-1"
          },
          "fixed_version": "0",
          "urgency": "unimportant"
        }
      }
    }
  }
}

比如,如果我的列表里有389-ds-base,希望能打印出这个包名。


问题原因与修复方案

原代码的问题

  • 转字符串搜索既低效又易出错:json.dumps(info)把JSON字典转成字符串后,用line in json_file匹配会因为文本行末尾的换行符(\n)导致匹配失败——比如文本里的nftables实际是nftables\n,而JSON字符串里的键是"nftables",完全不匹配。同时这种方式遍历整个字符串,数据量大时速度极慢。
  • 没利用JSON的结构化特性:加载后的info是Python字典,包名就是字典的顶级键,直接检查键是否存在才是正确做法。

修复后的代码

import json

def json_find():
    # 假设info已经通过json.load()或response.json()加载完成
    # 把包名转成集合,提升查找效率
    debian_packages = set(info.keys())
    
    with open("package_names.txt", "r") as f:
        for line in f:
            # 去除每行的换行符和首尾空白
            package = line.strip()
            # 直接检查包名是否在Debian包集合中
            if package in debian_packages:
                print(package)

json_find()

额外提示

  • 如果是从网络获取JSON数据,用requests.get(url).json()直接得到字典,无需手动处理字符串。
  • 用集合存储包名是因为集合的成员检查比字典略快,数据量越大优势越明显。

内容的提问来源于stack exchange,提问作者Kobotecnico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 09:15:44