Git仓库文件路径与文件名混淆处理技术需求问询
解决方案
核心思路
- 对每个文件夹/文件的名称部分(不含扩展名)生成固定哈希字符串,确保同一名称对应相同混淆值
- 完整保留文件扩展名
- 用字典缓存已处理的路径节点,避免重复计算,提升20万+文件的处理效率
- 输出混淆后的路径CSV,同时留存实际路径与混淆路径的映射关系
代码实现(Python)
import csv import hashlib from pathlib import Path def generate_hash(name): # 生成固定长度混淆字符串,可替换为SHA-1等其他哈希算法 return hashlib.md5(name.encode('utf-8')).hexdigest()[:8] # 取前8位足够区分不同名称 def obfuscate_path(original_path, cache): path_obj = Path(original_path) obfuscated_parts = [] # 逐层处理文件夹层级 current_path = Path('.') for part in path_obj.parent.parts: current_path /= part str_current = str(current_path) if str_current not in cache: cache[str_current] = generate_hash(part) obfuscated_parts.append(cache[str_current]) # 处理文件名,保留扩展名 filename = path_obj.name if '.' in filename: name_part, ext = filename.rsplit('.', 1) obfuscated_name = f"{generate_hash(name_part)}.{ext}" else: obfuscated_name = generate_hash(filename) obfuscated_parts.append(obfuscated_name) return '/'.join(obfuscated_parts) def process_csv(input_csv, output_obfuscated, output_mapping): cache = {} with open(input_csv, 'r', encoding='utf-8') as infile, \ open(output_obfuscated, 'w', newline='', encoding='utf-8') as outfile, \ open(output_mapping, 'w', newline='', encoding='utf-8') as mapfile: reader = csv.reader(infile) writer_obf = csv.writer(outfile) writer_map = csv.writer(mapfile) # 写入表头(假设原CSV第一列为路径) header = next(reader) writer_obf.writerow(header) writer_map.writerow(['original_path', 'obfuscated_path']) for row in reader: original_path = row[0] obfuscated = obfuscate_path(original_path, cache) writer_obf.writerow([obfuscated] + row[1:]) writer_map.writerow([original_path, obfuscated]) # 执行示例 if __name__ == '__main__': process_csv('original_paths.csv', 'obfuscated_paths.csv', 'path_mappings.csv')
关键说明
- 哈希稳定性:MD5哈希保证同一输入生成固定输出,严格满足"同一文件夹路径对应相同混淆映射"的要求
- 性能优化:
cache字典缓存已处理的文件夹路径,避免重复计算,大幅提升大文件量下的处理速度 - 扩展名保留:拆分文件名与扩展名,仅混淆名称部分,确保供应商能识别文件类型
- 映射追溯:单独生成映射CSV,方便内部随时关联实际路径与混淆路径
注意事项
- 若担心哈希冲突,可改用SHA-1并取前10位字符串,冲突概率可忽略
- 确保输入CSV路径格式统一,Windows风格
\可通过Path对象自动转换为Unix风格/ - 超大规模文件处理时,可分批次读取CSV,避免内存占用过高
内容的提问来源于stack exchange,提问作者hns
相关产品推荐
相关产品推荐

