You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Git仓库文件路径与文件名混淆处理技术需求问询

解决方案

核心思路

  • 对每个文件夹/文件的名称部分(不含扩展名)生成固定哈希字符串,确保同一名称对应相同混淆值
  • 完整保留文件扩展名
  • 用字典缓存已处理的路径节点,避免重复计算,提升20万+文件的处理效率
  • 输出混淆后的路径CSV,同时留存实际路径与混淆路径的映射关系

代码实现(Python)

import csv
import hashlib
from pathlib import Path

def generate_hash(name):
    # 生成固定长度混淆字符串,可替换为SHA-1等其他哈希算法
    return hashlib.md5(name.encode('utf-8')).hexdigest()[:8]  # 取前8位足够区分不同名称

def obfuscate_path(original_path, cache):
    path_obj = Path(original_path)
    obfuscated_parts = []
    
    # 逐层处理文件夹层级
    current_path = Path('.')
    for part in path_obj.parent.parts:
        current_path /= part
        str_current = str(current_path)
        if str_current not in cache:
            cache[str_current] = generate_hash(part)
        obfuscated_parts.append(cache[str_current])
    
    # 处理文件名,保留扩展名
    filename = path_obj.name
    if '.' in filename:
        name_part, ext = filename.rsplit('.', 1)
        obfuscated_name = f"{generate_hash(name_part)}.{ext}"
    else:
        obfuscated_name = generate_hash(filename)
    obfuscated_parts.append(obfuscated_name)
    
    return '/'.join(obfuscated_parts)

def process_csv(input_csv, output_obfuscated, output_mapping):
    cache = {}
    
    with open(input_csv, 'r', encoding='utf-8') as infile, \
         open(output_obfuscated, 'w', newline='', encoding='utf-8') as outfile, \
         open(output_mapping, 'w', newline='', encoding='utf-8') as mapfile:
        
        reader = csv.reader(infile)
        writer_obf = csv.writer(outfile)
        writer_map = csv.writer(mapfile)
        
        # 写入表头(假设原CSV第一列为路径)
        header = next(reader)
        writer_obf.writerow(header)
        writer_map.writerow(['original_path', 'obfuscated_path'])
        
        for row in reader:
            original_path = row[0]
            obfuscated = obfuscate_path(original_path, cache)
            writer_obf.writerow([obfuscated] + row[1:])
            writer_map.writerow([original_path, obfuscated])

# 执行示例
if __name__ == '__main__':
    process_csv('original_paths.csv', 'obfuscated_paths.csv', 'path_mappings.csv')

关键说明

  • 哈希稳定性:MD5哈希保证同一输入生成固定输出,严格满足"同一文件夹路径对应相同混淆映射"的要求
  • 性能优化:cache字典缓存已处理的文件夹路径,避免重复计算,大幅提升大文件量下的处理速度
  • 扩展名保留:拆分文件名与扩展名,仅混淆名称部分,确保供应商能识别文件类型
  • 映射追溯:单独生成映射CSV,方便内部随时关联实际路径与混淆路径

注意事项

  • 若担心哈希冲突,可改用SHA-1并取前10位字符串,冲突概率可忽略
  • 确保输入CSV路径格式统一,Windows风格\可通过Path对象自动转换为Unix风格/
  • 超大规模文件处理时,可分批次读取CSV,避免内存占用过高

内容的提问来源于stack exchange,提问作者hns

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 03:31:22