Azure Automation Runbook因内存不足挂起问题求助
问题描述
我通过PowerShell创建了Azure Automation账户,用于获取特定存储容器内文件的名称和大小,该代码通过Azure Data Factory的Webhook活动运行。处理中小型容器时代码运行正常,但处理包含大量文件的容器时,作业尝试3次后挂起。日志提示:
The runbook job failed due to sandbox running out of memory. Each Azure Automation sandbox is allocated 400 MB of memory. The job was attempted 3 times before it was suspended.
请问如何解决该问题,是否可以增加内存?
附原PowerShell代码:
#define parameters param ( [Parameter (Mandatory = $false)] [object] $WebHookData, [string]$StorageAccountName, [string]$StorageAccountKey ) $Parameters = (ConvertFrom-Json -InputObject $WebHookData.RequestBody) <#If ($Parameters.callBackUri) { $callBackUri = $Parameters.callBackUri }#> $containerName = $Parameters.containerName "->"+$StorageAccountName "->"+$StorageAccountKey $connectionName = "AzureRunAsConnection" try { # Get the connection "AzureRunAsConnection " $servicePrincipalConnection=Get-AutomationConnection -Name $connectionName "Logging in to Azure..." Connect-AzAccount ` -ServicePrincipal ` -TenantId $servicePrincipalConnection.TenantId ` -ApplicationId $servicePrincipalConnection.ApplicationId ` -CertificateThumbprint $servicePrincipalConnection.CertificateThumbprint } catch { if (!$servicePrincipalConnection) { $ErrorMessage = "Connection $connectionName not found." throw $ErrorMessage } else{ Write-Error -Message $_.Exception throw $_.Exception } } #storage account $StorageAccountName = $StorageAccountName #storage key $StorageAccountKey = $StorageAccountKey #Container name - change if different $containerName = $containerName #get blob context $Ctx = New-AzStorageContext $StorageAccountName -StorageAccountKey $StorageAccountKey # get a list of all of the blobs in the container $listOfBlobs = Get-AzStorageBlob -Container $containerName -Context $Ctx # zero out our total $length = 0 # this loops through the list of blobs and retrieves the length for each blob # and adds it to the total $listOfBlobs | ForEach-Object {$length = $length + $_.Length} # output the blobs and their sizes and the total Write-Host "List of Blobs and their size (length)" Write-Host " " $select = $listOfBlobs | Select-Object -Property @{Name='ContainerName';Expression={$containerName}}, Name, @{name="Size";expression={$($_.Length)}}, LastModified #$listOfBlobs | select Name, Length, @{Name='ContainerName';Expression={$containerName}} Write-Host " " Write-Host "Total Length = " $length #define location and Export to CSV file $SourceLocation = Get-Location $select | Export-Csv $SourceLocation'File-size/File-size-'$containerName'.csv' -NoTypeInformation -Force -Encoding UTF8 $Context = New-AzureStorageContext -StorageAccountName $StorageAccountName -StorageAccountKey $StorageAccountKey Set-AzureStorageBlobContent -Context $Context -Container "Name" -File $SourceLocation"File-size/File-size-$containerName.csv" -Blob "File-Size/File-size-$containerName.csv" -Force
解决方案
是否可以增加内存?
Azure Automation沙箱的400MB内存是固定配额,无法直接调整。必须通过优化代码逻辑来降低内存占用。
代码优化方案
原代码一次性加载所有Blob对象到内存,当容器内文件数量极大时,会导致内存耗尽。优化核心是分批处理Blob,避免一次性加载全部数据:
优化后的PowerShell代码
param ( [Parameter (Mandatory = $false)] [object] $WebHookData, [string]$StorageAccountName, [string]$StorageAccountKey ) $Parameters = (ConvertFrom-Json -InputObject $WebHookData.RequestBody) $containerName = $Parameters.containerName $connectionName = "AzureRunAsConnection" try { $servicePrincipalConnection = Get-AutomationConnection -Name $connectionName Connect-AzAccount ` -ServicePrincipal ` -TenantId $servicePrincipalConnection.TenantId ` -ApplicationId $servicePrincipalConnection.ApplicationId ` -CertificateThumbprint $servicePrincipalConnection.CertificateThumbprint } catch { if (!$servicePrincipalConnection) { throw "Connection $connectionName not found." } else { Write-Error -Message $_.Exception throw $_.Exception } } $Ctx = New-AzStorageContext $StorageAccountName -StorageAccountKey $StorageAccountKey $length = 0 $csvFilePath = Join-Path (Get-Location) "File-size/File-size-$containerName.csv" # 创建CSV文件表头 @{ ContainerName = $containerName Name = "" Size = 0 LastModified = $null } | Select-Object ContainerName, Name, Size, LastModified | Export-Csv -Path $csvFilePath -NoTypeInformation -Force -Encoding UTF8 # 分批获取Blob,每次处理1000条 $continuationToken = $null do { $blobBatch = Get-AzStorageBlob -Container $containerName -Context $Ctx -MaxCount 1000 -ContinuationToken $continuationToken $continuationToken = $blobBatch.ContinuationToken # 处理当前批次的Blob $batchData = $blobBatch | ForEach-Object { $length += $_.Length [PSCustomObject]@{ ContainerName = $containerName Name = $_.Name Size = $_.Length LastModified = $_.LastModified } } # 追加到CSV文件 $batchData | Export-Csv -Path $csvFilePath -NoTypeInformation -Append -Force -Encoding UTF8 # 释放当前批次内存 Remove-Variable blobBatch, batchData -ErrorAction SilentlyContinue [GC]::Collect() } while ($continuationToken -ne $null) Write-Host "Total Length = " $length # 上传CSV到存储容器 Set-AzureStorageBlobContent -Context $Ctx -Container "Name" -File $csvFilePath -Blob "File-Size/File-size-$containerName.csv" -Force # 清理临时文件 Remove-Item $csvFilePath -Force -ErrorAction SilentlyContinue
优化点说明
- 分批获取Blob:使用
-MaxCount和-ContinuationToken参数,每次只加载1000条Blob数据,避免一次性加载全部数据到内存。 - 增量写入CSV:先创建表头,再分批追加数据,避免一次性保存所有Blob对象到内存中的数组。
- 主动释放内存:处理完每批次后,删除变量并触发垃圾回收,及时释放内存。
- 减少重复操作:复用同一个存储上下文,避免重复创建
New-AzureStorageContext实例。
其他可选优化建议
- 使用Azure Functions替代:如果容器内文件数量极大,Azure Functions提供更高的内存配额(最高14GB),适合处理大规模数据。
- 筛选Blob:如果只需要特定前缀或类型的文件,在
Get-AzStorageBlob中添加-Prefix参数减少数据量。
内容的提问来源于stack exchange,提问作者coding
相关产品推荐
相关产品推荐

