You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大体积TXT文件按ID分组校验条件及高效清理的技术求助

问题描述

我有一个1300MB的超大TXT文件,需要实现两个功能:

  • 每行开头包含唯一ID,需按ID分组校验条件,统计每组中符合条件的行数;
  • 脚本执行完成后移除所有符合条件的行,以便使用新条件重复执行来逐步缩窄文件范围。

经过几轮操作后需得到适用于所有行的条件集,但当前PowerShell脚本运行极慢(单次循环需数小时),现有代码如下:

foreach ($item in $liste)
{
    
    # Check Conditions
    if ( ($item -like "*XXX*") -and ($item -like "*YYY*") -and ($item -notlike "*ZZZ*")) { 
        
     # Add a line to a document to see which lines match condition                    
        Add-Content "C:\Desktop\it_seems_to_match.txt" "$item"
        
    # Retrieve the unique ID from the line and feed array.                
        $array += $item.Split("/")[1]

    # Remove the line from final document
        $liste = $liste -replace $item, ""         
            
    }

                              
} 
# Pipe the "new cleaned" list somewhere
    $liste | Set-Content -Path "C:\NewListToWorkWith.txt"
# Show me the counts
    $array | group | % { $h = @{} } { $h[$_.Name] = $_.Count } { $h } | Out-File "C:\Desktop\count.txt"

示例行:

images/STRINGA/2XXXXXXXX_rTTTTw_GGGG1_Top_MMM1_YY02_ZZZ30_AAAA5.jpg images/STRINGA/3XXXXXXXX_rTTTTw_GGGG1_Top_MMM1_YY02_ZZZ30_AAAA5.jpg images/STRINGB/4XXXXXXXX_rTTTTw_GGGG1_Top_MMM1_YY02_ZZZ30_AAAA5.jpg images/STRINGB/5XXXXXXXX_rTTTTw_GGGG1_Top_MMM1_YY02_ZZZ30_AAAA5.jpg images/STRINGC/5XXXXXXXX_rTTTTw_GGGG1_Top_MMM1_YY02_ZZZ30_AAAA5.jpg


优化方案

原脚本性能瓶颈在于:频繁IO操作、数组扩容开销、O(n²)的替换逻辑、全量加载大文件占用内存。以下是针对性优化实现:

核心优化思路

  • 逐行读取文件,避免一次性加载全部内容到内存;
  • 使用文件流批量写入,减少IO打开/关闭次数;
  • 用哈希表实时统计ID计数,避免后期分组操作;
  • 分离匹配与非匹配行,直接写入目标文件,无需替换操作。

优化后的代码

# 定义文件路径
$inputPath = "C:\YourOriginalFile.txt"
$matchOutputPath = "C:\Desktop\it_seems_to_match.txt"
$cleanOutputPath = "C:\NewListToWorkWith.txt"
$countOutputPath = "C:\Desktop\count.txt"

# 初始化哈希表用于ID计数
$idCounts = @{}

# 创建文件流写入器,提升IO效率
$matchWriter = [System.IO.StreamWriter]::new($matchOutputPath, $false)
$cleanWriter = [System.IO.StreamWriter]::new($cleanOutputPath, $false)

try {
    # 逐行读取大文件,降低内存占用
    foreach ($line in [System.IO.File]::ReadLines($inputPath)) {
        # 条件校验
        $isMatch = ($line -like "*XXX*") -and ($line -like "*YYY*") -and ($line -notlike "*ZZZ*")
        
        if ($isMatch) {
            $matchWriter.WriteLine($line)
            # 提取ID并更新计数
            $id = $line.Split("/")[1]
            $idCounts[$id] = ($idCounts[$id] ?? 0) + 1
        } else {
            $cleanWriter.WriteLine($line)
        }
    }
} finally {
    # 确保资源释放
    $matchWriter.Dispose()
    $cleanWriter.Dispose()
}

# 输出统计结果
$idCounts.GetEnumerator() | ForEach-Object { "$($_.Key): $($_.Value)" } | Set-Content $countOutputPath

额外优化建议

  • 正则替代通配符:复杂条件下用[regex]::IsMatch($line, 'XXX.*YYY(?!.*ZZZ)')替代多个-like,匹配效率更高;
  • 指定编码:若原文件为ASCII格式,创建StreamWriter时添加编码参数[System.Text.Encoding]::ASCII,进一步提升IO速度;
  • 递进式条件合并:如果多轮条件是递进筛选,可在一次遍历中应用所有条件,减少文件读取次数;
  • 内存监控:逐行读取模式下,内存占用仅为单行列的大小,避免大文件导致的内存溢出或GC卡顿。

内容的提问来源于stack exchange,提问作者Julian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 05:10:27