You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

140K个DL迁移批次划分:基于CSV关联分组方案需求

分发组(DL)迁移批次划分方案

需求背景

现有140K个分发组需要迁移,需从包含所有DL的CSV文件中划分可追踪的迁移批次:

  • CSV规则:若某DL标记为某个Set,其关联列中包含的其他DL必须归为同一Set
  • 关联关系支持互相引用、多层嵌套(例:DL1与DL2互相引用同属Set-1;DL10关联DL4/5/6则均属Set-4)
  • 现有PowerShell代码仅支持实时查询AD,存在循环重复等缺陷,无法直接处理CSV文件,需替换为直接处理CSV的方案

现有问题代码(PowerShell)

$Groups = New-Object -TypeName "System.Collections.ArrayList"
$Groups = [System.Collections.ArrayList]@()
$dump = $Groups.Add("$($GroupName)")
$final = New-Object -TypeName "System.Collections.ArrayList"
$final = [System.Collections.ArrayList]@()
for($i=0; $i -lt $Groups.count; $i++){
$group = $Groups[$i]
$members = Get-DistributionGroupMember -Identity $Group
foreach($member in $members){
if($member.RecipientType -like "*Group"){
$dump = $Groups.Add("$($member.Name)")
}
else{
$dump = $final.Add("$($member.Name)")
}
}
}
$final = $final | Sort -Unique
$final | Out-File -FilePath ".\MemberList_$($GroupName).txt"
$path = pwd
Write-Host "The member list can be found at $($path)\MemberList_$($GroupName).txt" -ForegroundColor Green

解决方案

方案1:PowerShell直接处理CSV

核心用**并查集(Union-Find)**算法处理关联关系,遍历CSV构建组间关联,最终划分批次:

# 加载CSV文件(假设列名为:GroupName, AssociatedGroups,AssociatedGroups用逗号分隔多个DL)
$dlData = Import-Csv -Path ".\DL_List.csv" -Delimiter ","

# 初始化并查集字典:键为DL名称,值为其父节点
$parent = @{}
$dlData | ForEach-Object {
    $parent[$_.GroupName] = $_.GroupName
}

# 查找根节点(带路径压缩)
function Find-Root {
    param($node)
    if ($parent[$node] -ne $node) {
        $parent[$node] = Find-Root -Node $parent[$node]
    }
    return $parent[$node]
}

# 合并两个节点所属集合
function Merge-Sets {
    param($node1, $node2)
    $root1 = Find-Root -Node $node1
    $root2 = Find-Root -Node $node2
    if ($root1 -ne $root2) {
        $parent[$root2] = $root1
    }
}

# 遍历所有DL,处理关联关系
$dlData | ForEach-Object {
    $currentDL = $_.GroupName
    # 拆分关联的DL列表
    $associatedDLs = $_.AssociatedGroups -split "," | ForEach-Object { $_.Trim() } | Where-Object { $_ -ne "" }
    foreach ($dl in $associatedDLs) {
        if ($parent.ContainsKey($dl)) { # 确保关联的DL存在于列表中
            Merge-Sets -Node1 $currentDL -Node2 $dl
        }
    }
}

# 生成批次结果:按根节点分组,分配Set编号
$setCounter = 1
$batchResults = @()
$groupedByRoot = $parent.GetEnumerator() | Group-Object -Property Value

foreach ($group in $groupedByRoot) {
    $setName = "Set-$setCounter"
    foreach ($entry in $group.Group) {
        $batchResults += [PSCustomObject]@{
            GroupName = $entry.Key
            BatchSet  = $setName
            RootGroup = $group.Name
        }
    }
    $setCounter++
}

# 导出结果到CSV
$batchResults | Export-Csv -Path ".\DL_Migration_Batches.csv" -NoTypeInformation -Encoding UTF8
Write-Host "迁移批次划分完成,结果已导出到DL_Migration_Batches.csv" -ForegroundColor Green

方案2:Python直接处理CSV

同样基于并查集算法,适合处理超大规模数据(140K量级无压力):

import csv
from collections import defaultdict

class UnionFind:
    def __init__(self):
        self.parent = {}
    
    def find(self, node):
        if self.parent[node] != node:
            self.parent[node] = self.find(self.parent[node])
        return self.parent[node]
    
    def union(self, node1, node2):
        root1 = self.find(node1)
        root2 = self.find(node2)
        if root1 != root2:
            self.parent[root2] = root1

# 初始化并查集
uf = UnionFind()
dl_list = []

# 读取CSV文件(假设列名:GroupName, AssociatedGroups)
with open('DL_List.csv', 'r', encoding='utf-8') as f:
    reader = csv.DictReader(f)
    for row in reader:
        dl_name = row['GroupName'].strip()
        uf.parent[dl_name] = dl_name
        dl_list.append((dl_name, row['AssociatedGroups'].strip()))

# 处理关联关系
for dl_name, associated in dl_list:
    if not associated:
        continue
    associated_dls = [d.strip() for d in associated.split(',') if d.strip()]
    for adl in associated_dls:
        if adl in uf.parent:
            uf.union(dl_name, adl)

# 按根节点分组,生成批次
root_to_dls = defaultdict(list)
for dl in uf.parent:
    root = uf.find(dl)
    root_to_dls[root].append(dl)

# 生成结果并导出
set_counter = 1
result_rows = []
for root, dls in root_to_dls.items():
    set_name = f"Set-{set_counter}"
    for dl in dls:
        result_rows.append({
            'GroupName': dl,
            'BatchSet': set_name,
            'RootGroup': root
        })
    set_counter += 1

with open('DL_Migration_Batches.csv', 'w', encoding='utf-8', newline='') as f:
    writer = csv.DictWriter(f, fieldnames=['GroupName', 'BatchSet', 'RootGroup'])
    writer.writeheader()
    writer.writerows(result_rows)

print("迁移批次划分完成,结果已保存到DL_Migration_Batches.csv")

备注

  • Excel公式不推荐用于140K量级数据,公式嵌套复杂且性能极差,容易崩溃
  • 两个方案均支持互相引用、多层嵌套的关联关系,确保关联的DL归为同一批次
  • 处理前需确保CSV文件中AssociatedGroups列的DL名称与GroupName列完全匹配(无大小写、空格差异)

内容的提问来源于stack exchange,提问作者Jeetcu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 09:45:50