You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

需将递归反向列文件的Bash脚本转换为PowerShell,处理海量JSON文件

Let's break down how to convert your Bash recursive file import script to PowerShell, plus some critical optimizations for handling your massive scale (700k folders, 87M JSON files) since WSL isn't playing nice with your mounted Google Drive.


First, here's a direct translation of your Bash import_from_start function. This will recursively find all .json files starting from your top-level parent folder and process them:

function Import-FromStart {
    param(
        [Parameter(Mandatory=$true)]
        [string]$RootPath
    )

    Write-Host "Starting import from front"
    # Recursively find all JSON files (equivalent to Bash's `find "$1" -name '*.json'`)
    Get-ChildItem -Path $RootPath -Filter '*.json' -Recurse -File | ForEach-Object {
        # Replace this with your actual import logic
        Write-Verbose "Processing file: $($_.FullName)"
        # Example: $jsonData = Get-Content $_.FullName -Raw | ConvertFrom-Json
        # Then import $jsonData to your target system/database
    }
}

# Call the function with your top-level parent folder path
Import-FromStart -RootPath "G:\YourMountedGoogleDrive" -Verbose

Optimizations for Your Massive File Set

87 million files is no joke—basic recursion will be way too slow. Here are tailored optimizations to handle this scale:

1. Process Deepest Folders First (Reverse Recursion)

If your Bash script intended to process nested folders from the deepest level upward, we can calculate folder depth and sort accordingly:

function Import-FromDeepestFirst {
    param(
        [Parameter(Mandatory=$true)]
        [string]$RootPath
    )

    Write-Host "Starting import from deepest folders first"
    
    # Get all folders and calculate their depth relative to the root
    $folders = Get-ChildItem -Path $RootPath -Directory -Recurse | Select-Object FullName, @{
        Name = 'Depth'
        Expression = { ($_.FullName.Split('\').Count) - ($RootPath.Split('\').Count) }
    }

    # Sort folders from deepest to shallowest
    $sortedFolders = $folders | Sort-Object -Property Depth -Descending

    # Process each folder's JSON files
    foreach ($folder in $sortedFolders) {
        Write-Verbose "Processing folder: $($folder.FullName) (depth: $($folder.Depth))"
        Get-ChildItem -Path $folder.FullName -Filter '*.json' -File | ForEach-Object {
            # Your import logic here
        }
    }

    # Finally process JSON files in the root folder
    Get-ChildItem -Path $RootPath -Filter '*.json' -File | ForEach-Object {
        # Your import logic here
    }
}

2. Parallel Processing (Huge Speed Boost)

PowerShell 7+ supports parallel execution, which is a must for handling millions of files. Adjust the throttle limit based on your CPU and disk capabilities:

function Import-Parallel {
    param(
        [Parameter(Mandatory=$true)]
        [string]$RootPath,
        [int]$ThrottleLimit = 8 # Tune this based on your system
    )

    Write-Host "Starting parallel import"
    
    # Collect all JSON file paths first (avoids repeated folder traversal)
    $jsonFiles = Get-ChildItem -Path $RootPath -Filter '*.json' -Recurse -File | Select-Object -ExpandProperty FullName

    # Process files in parallel
    $jsonFiles | ForEach-Object -Parallel {
        $filePath = $_
        try {
            # Your import logic here
            $jsonContent = Get-Content -Path $filePath -Raw -ErrorAction Stop | ConvertFrom-Json -ErrorAction Stop
            # Example: Invoke-SqlCmd -ServerInstance "YourDBServer" -Database "YourDB" -Query "INSERT ..." -Variable $jsonContent
            Write-Verbose "Processed: $filePath"
        }
        catch {
            # Log errors instead of stopping the entire process
            "[$(Get-Date)] Failed to process $filePath : $_" | Out-File -Path "import_errors.log" -Append
        }
    } -ThrottleLimit $ThrottleLimit -Verbose
}

3. Batched Processing (Reduce Memory Overhead)

Collecting 87 million file paths in memory will drain resources. Instead, process files in batches:

function Import-Batched {
    param(
        [Parameter(Mandatory=$true)]
        [string]$RootPath,
        [int]$BatchSize = 1000,
        [int]$ThrottleLimit = 8
    )

    Write-Host "Starting batched parallel import"
    
    # Helper function to split objects into batches
    function Batch-Object {
        param(
            [Parameter(ValueFromPipeline=$true)]
            [object]$InputObject,
            [int]$Size = 100
        )
        begin { $batch = @() }
        process {
            $batch += $InputObject
            if ($batch.Count -ge $Size) {
                $batch
                $batch = @()
            }
        }
        end { if ($batch.Count -gt 0) { $batch } }
    }

    # Traverse folders, batch files, and process in parallel
    Get-ChildItem -Path $RootPath -Filter '*.json' -Recurse -File | 
        Select-Object -ExpandProperty FullName |
        Batch-Object -Size $BatchSize |
        ForEach-Object {
            $_ | ForEach-Object -Parallel {
                $filePath = $_
                # Your import logic here (with error handling)
            } -ThrottleLimit $ThrottleLimit
        }
}

Critical Notes for Google Drive Mounts
  • Performance Limits: PowerShell directly accessing a mounted Google Drive may still have latency. If possible, sync chunks of files to your local disk first before processing.
  • PowerShell Version: Use PowerShell 7+ (not Windows PowerShell 5.1) for better parallelism and performance.
  • Database Bulk Inserts: If importing to a database, use bulk insert commands instead of single-row inserts—this will cut down on connection overhead drastically.

内容的提问来源于stack exchange,提问作者Gabriel Fair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:04:39