需将递归反向列文件的Bash脚本转换为PowerShell,处理海量JSON文件
Let's break down how to convert your Bash recursive file import script to PowerShell, plus some critical optimizations for handling your massive scale (700k folders, 87M JSON files) since WSL isn't playing nice with your mounted Google Drive.
First, here's a direct translation of your Bash import_from_start function. This will recursively find all .json files starting from your top-level parent folder and process them:
function Import-FromStart { param( [Parameter(Mandatory=$true)] [string]$RootPath ) Write-Host "Starting import from front" # Recursively find all JSON files (equivalent to Bash's `find "$1" -name '*.json'`) Get-ChildItem -Path $RootPath -Filter '*.json' -Recurse -File | ForEach-Object { # Replace this with your actual import logic Write-Verbose "Processing file: $($_.FullName)" # Example: $jsonData = Get-Content $_.FullName -Raw | ConvertFrom-Json # Then import $jsonData to your target system/database } } # Call the function with your top-level parent folder path Import-FromStart -RootPath "G:\YourMountedGoogleDrive" -Verbose
87 million files is no joke—basic recursion will be way too slow. Here are tailored optimizations to handle this scale:
1. Process Deepest Folders First (Reverse Recursion)
If your Bash script intended to process nested folders from the deepest level upward, we can calculate folder depth and sort accordingly:
function Import-FromDeepestFirst { param( [Parameter(Mandatory=$true)] [string]$RootPath ) Write-Host "Starting import from deepest folders first" # Get all folders and calculate their depth relative to the root $folders = Get-ChildItem -Path $RootPath -Directory -Recurse | Select-Object FullName, @{ Name = 'Depth' Expression = { ($_.FullName.Split('\').Count) - ($RootPath.Split('\').Count) } } # Sort folders from deepest to shallowest $sortedFolders = $folders | Sort-Object -Property Depth -Descending # Process each folder's JSON files foreach ($folder in $sortedFolders) { Write-Verbose "Processing folder: $($folder.FullName) (depth: $($folder.Depth))" Get-ChildItem -Path $folder.FullName -Filter '*.json' -File | ForEach-Object { # Your import logic here } } # Finally process JSON files in the root folder Get-ChildItem -Path $RootPath -Filter '*.json' -File | ForEach-Object { # Your import logic here } }
2. Parallel Processing (Huge Speed Boost)
PowerShell 7+ supports parallel execution, which is a must for handling millions of files. Adjust the throttle limit based on your CPU and disk capabilities:
function Import-Parallel { param( [Parameter(Mandatory=$true)] [string]$RootPath, [int]$ThrottleLimit = 8 # Tune this based on your system ) Write-Host "Starting parallel import" # Collect all JSON file paths first (avoids repeated folder traversal) $jsonFiles = Get-ChildItem -Path $RootPath -Filter '*.json' -Recurse -File | Select-Object -ExpandProperty FullName # Process files in parallel $jsonFiles | ForEach-Object -Parallel { $filePath = $_ try { # Your import logic here $jsonContent = Get-Content -Path $filePath -Raw -ErrorAction Stop | ConvertFrom-Json -ErrorAction Stop # Example: Invoke-SqlCmd -ServerInstance "YourDBServer" -Database "YourDB" -Query "INSERT ..." -Variable $jsonContent Write-Verbose "Processed: $filePath" } catch { # Log errors instead of stopping the entire process "[$(Get-Date)] Failed to process $filePath : $_" | Out-File -Path "import_errors.log" -Append } } -ThrottleLimit $ThrottleLimit -Verbose }
3. Batched Processing (Reduce Memory Overhead)
Collecting 87 million file paths in memory will drain resources. Instead, process files in batches:
function Import-Batched { param( [Parameter(Mandatory=$true)] [string]$RootPath, [int]$BatchSize = 1000, [int]$ThrottleLimit = 8 ) Write-Host "Starting batched parallel import" # Helper function to split objects into batches function Batch-Object { param( [Parameter(ValueFromPipeline=$true)] [object]$InputObject, [int]$Size = 100 ) begin { $batch = @() } process { $batch += $InputObject if ($batch.Count -ge $Size) { $batch $batch = @() } } end { if ($batch.Count -gt 0) { $batch } } } # Traverse folders, batch files, and process in parallel Get-ChildItem -Path $RootPath -Filter '*.json' -Recurse -File | Select-Object -ExpandProperty FullName | Batch-Object -Size $BatchSize | ForEach-Object { $_ | ForEach-Object -Parallel { $filePath = $_ # Your import logic here (with error handling) } -ThrottleLimit $ThrottleLimit } }
- Performance Limits: PowerShell directly accessing a mounted Google Drive may still have latency. If possible, sync chunks of files to your local disk first before processing.
- PowerShell Version: Use PowerShell 7+ (not Windows PowerShell 5.1) for better parallelism and performance.
- Database Bulk Inserts: If importing to a database, use bulk insert commands instead of single-row inserts—this will cut down on connection overhead drastically.
内容的提问来源于stack exchange,提问作者Gabriel Fair

