Python处理复杂数据结构:随机访问场景下的最优方案咨询
打印机映射记录的存储与查询最优方案
方案思路
用dataclass存储单条记录保证结构清晰,同时构建多维度索引字典来满足所有查询需求,兼顾可读性与查询效率。
1. 定义数据类存储单条记录
用dataclass封装每条映射的四个字段,结构直观且支持不可变特性(避免意外修改):
from dataclasses import dataclass from typing import List, Dict, Set @dataclass(frozen=True) class PrinterMapping: user: str company: str doc_type: str printer: str
2. 加载文件并构建查询索引
读取文件时,将每行解析为PrinterMapping实例,同时构建4个索引字典,分别对应你的查询需求:
def load_mappings(file_path: str) -> tuple[List[PrinterMapping], Dict[str, Set[PrinterMapping]], Dict[str, Set[str]], Dict[str, Set[str]], Dict[tuple[str, str, str], PrinterMapping]]: mappings = [] # 索引1:用户 -> 该用户的所有映射记录 user_index = {} # 索引2:打印机 -> 使用该打印机的用户集合 printer_to_users = {} # 索引3:文档类型 -> 拥有该类型的用户集合 doc_type_to_users = {} # 索引4:(用户, 打印机, 文档类型) -> 对应映射记录(用于快速验证) triple_index = {} with open(file_path, 'r') as f: for line in f: line = line.strip() if not line: continue user, company, doc_type, printer = line.split(':') mapping = PrinterMapping(user, company, doc_type, printer) mappings.append(mapping) # 更新用户索引 user_index.setdefault(user, set()).add(mapping) # 更新打印机-用户索引 printer_to_users.setdefault(printer, set()).add(user) # 更新文档类型-用户索引 doc_type_to_users.setdefault(doc_type, set()).add(user) # 更新复合验证索引 triple_index[(user, printer, doc_type)] = mapping return mappings, user_index, printer_to_users, doc_type_to_users, triple_index
3. 实现查询函数
基于预构建的索引,直接实现各查询需求,所有查询均为O(1)或O(k)(k为目标分组内的记录数),效率极高:
查询1:列出特定用户的所有打印机
def get_user_printers(user: str, user_index: Dict[str, Set[PrinterMapping]]) -> Set[str]: return {m.printer for m in user_index.get(user, set())}
查询2:列出使用特定打印机的所有用户
def get_printer_users(printer: str, printer_to_users: Dict[str, Set[str]]) -> Set[str]: return printer_to_users.get(printer, set())
查询3:列出拥有特定文档类型的所有用户
def get_doc_type_users(doc_type: str, doc_type_to_users: Dict[str, Set[str]]) -> Set[str]: return doc_type_to_users.get(doc_type, set())
查询4:验证用户是否存在指定打印机与文档类型的映射
def validate_mapping(user: str, printer: str, doc_type: str, triple_index: Dict[tuple[str, str, str], PrinterMapping]) -> bool: return (user, printer, doc_type) in triple_index
方案优势
- 可读性强:
dataclass让每条记录的结构一目了然,比嵌套字典更易维护 - 查询高效:预构建索引避免了全表遍历,所有常用查询都能快速响应
- 扩展性好:后续新增查询需求(如按公司查询),只需新增对应索引即可
内容的提问来源于stack exchange,提问作者WJ.Lesster
相关产品推荐
相关产品推荐

