You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历多字典进行关键词搜索比对的效率优化咨询

效率优化方案:YAML数据存在性检查

当前的遍历方式绝对不是最优效率的实现方式——每次针对grant里的一个数据点就去遍历users/groups/apps文件夹,会产生大量重复的文件IO操作,尤其是当文件数量超过50个时,这种重复IO会严重拖慢整体处理速度。

最优实现思路:预加载+哈希索引

核心是把需要查询的基准数据(users/groups/apps里的目标键值)一次性加载到内存并构建快速查询结构,避免重复读取文件:

  • 第一步:预加载并构建索引
    先遍历users、groups、apps三个文件夹下的所有YAML文件,把每个文件里需要验证的目标键值(比如用户ID、组名、应用ID)提取出来,分别存入对应的哈希集合(如Python的set)。集合的查询时间复杂度是O(1),远快于遍历文件的O(n)。

    举个Python示例(用PyYAML库):

    import yaml
    import os
    
    def build_index(folder_path, target_key):
        index_set = set()
        for filename in os.listdir(folder_path):
            if filename.endswith('.yaml'):
                with open(os.path.join(folder_path, filename), 'r') as f:
                    data = yaml.safe_load(f)
                    # 兼容单字典或字典列表格式的YAML文件
                    if isinstance(data, list):
                        for item in data:
                            if target_key in item:
                                index_set.add(item[target_key])
                    elif isinstance(data, dict):
                        if target_key in data:
                            index_set.add(data[target_key])
        return index_set
    
    # 预构建三个索引集合
    user_ids = build_index('users', 'user_id')
    group_names = build_index('groups', 'group_name')
    app_ids = build_index('apps', 'app_id')
    
  • 第二步:批量处理grant数据
    遍历grants文件夹下的所有YAML文件,对每条grant数据直接通过集合查询验证键值是否存在,无需再去遍历users等文件夹:

    def process_grants(grants_folder):
        for filename in os.listdir(grants_folder):
            if filename.endswith('.yaml'):
                with open(os.path.join(grants_folder, filename), 'r') as f:
                    grants = yaml.safe_load(f)
                    for grant in grants:
                        # 验证示例
                        is_valid_user = grant.get('user_id') in user_ids
                        is_valid_group = grant.get('group_name') in group_names
                        is_valid_app = grant.get('app_id') in app_ids
                        # 后续处理逻辑
                        print(f"Grant {grant.get('id')} 验证结果:用户{is_valid_user},组{is_valid_group},应用{is_valid_app}")
    
    process_grants('grants')
    

额外优化建议

  • 如果users/groups/apps的文件量极大,可采用并行预加载(比如用Python的concurrent.futures),加快索引构建速度,但要注意YAML解析的线程安全问题。
  • 若YAML文件过大,可考虑流式解析(比如PyYAML的yaml.parse),避免一次性加载整个文件占用过多内存。

内容的提问来源于stack exchange,提问作者Barry Keegan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:20:05