如何使用jq流解析Amazon RDS实例信息并优化内存占用?
解决Amazon RDS实例数据解析的内存优化问题
我在解析Amazon RDS实例数据时,想要生成一个以instanceType为键、内存字节数为值的对象,但原表达式运行时内存占用接近1GB。尝试用jq的--stream模式优化后,目前只能分行输出键和对应值,需要进一步调整实现目标。
原高内存占用表达式
这个表达式能实现需求,但内存占用过高:
.products | to_entries | map(.value.attributes | select(.instanceType != null) | {(.instanceType): ((.memory | split(" ") | .[0] | tonumber) * 1024 * 1024 * 1024)}) | add
当前stream模式的尝试
由于在线环境不支持--stream参数,用(. | tostream)替代,当前表达式如下:
def isMemory: .[0][3] == "memory"; def isInstanceType: .[0][3] == "instanceType"; (. | tostream) | select(isMemory or isInstanceType) | .[1]
输出结果为分行的键值对:
"db.r5.24xlarge" "768 GiB" "db.r4.large" "15.25 GiB"
简化版测试JSON
{ "formatVersion" : "v1.0", "disclaimer" : "This pricing list is for informational purposes only. All prices are subject to the additional terms included in the pricing pages on http://aws.amazon.com. All Free Tier prices are also subject to the terms included at https://aws.amazon.com/free/", "offerCode" : "AmazonRDS", "version" : "20230328234721", "publicationDate" : "2023-03-28T23:47:21Z", "products" : { "BHYABS232JP4AGQY" : { "sku" : "BHYABS232JP4AGQY", "productFamily" : "Database Instance", "attributes" : { "servicecode" : "AmazonRDS", "location" : "US East (Ohio)", "locationType" : "AWS Region", "instanceType" : "db.r5.24xlarge", "currentGeneration" : "Yes", "instanceFamily" : "Memory optimized", "vcpu" : "96", "physicalProcessor" : "Intel Xeon Platinum 8175", "clockSpeed" : "Up to 3.1 GHz", "memory" : "768 GiB", "storage" : "EBS Only", "networkPerformance" : "25 Gigabit", "processorArchitecture" : "64-bit", "engineCode" : "18", "databaseEngine" : "MariaDB", "licenseModel" : "No license required", "deploymentOption" : "Single-AZ", "usagetype" : "USE2-InstanceUsage:db.r5.24xl", "operation" : "CreateDBInstance:0018", "dedicatedEbsThroughput" : "14000 Mbps", "enhancedNetworkingSupported" : "Yes", "instanceTypeFamily" : "R5", "normalizationSizeFactor" : "192", "regionCode" : "us-east-2", "servicename" : "Amazon Relational Database Service" } }, "D8GBHQEK73G5ADCK" : { "sku" : "D8GBHQEK73G5ADCK", "productFamily" : "Database Instance", "attributes" : { "servicecode" : "AmazonRDS", "location" : "Asia Pacific (Tokyo)", "locationType" : "AWS Region", "instanceType" : "db.r4.large", "currentGeneration" : "No", "instanceFamily" : "Memory optimized", "vcpu" : "2", "physicalProcessor" : "Intel Xeon E5-2686 v4 (Broadwell)", "clockSpeed" : "2.3 GHz", "memory" : "15.25 GiB", "storage" : "EBS Only", "networkPerformance" : "Up to 10 Gigabit", "processorArchitecture" : "64-bit", "engineCode" : "2", "databaseEngine" : "MySQL", "licenseModel" : "No license required", "deploymentOption" : "Multi-AZ", "usagetype" : "APN1-Multi-AZUsage:db.r4.large", "operation" : "CreateDBInstance:0002", "dedicatedEbsThroughput" : "400 Mbps", "enhancedNetworkingSupported" : "Yes", "instanceTypeFamily" : "R4", "normalizationSizeFactor" : "8", "processorFeatures" : "Intel AVX, Intel AVX2, Intel Turbo", "regionCode" : "ap-northeast-1", "servicename" : "Amazon Relational Database Service" } } } }
优化后的stream模式表达式
要在stream模式下构建目标对象,需要跟踪每个产品的instanceType和memory字段,当收集到一对完整的键值后,转换内存单位并添加到结果对象中。表达式如下:
# 跟踪当前实例类型 reduce (. | tostream) as $item ({}; if $item[0][0] == "products" and $item[0][2] == "attributes" then if $item[0][3] == "instanceType" then . + {currentType: $item[1]} elif $item[0][3] == "memory" and .currentType != null then # 转换内存为字节数 let memoryBytes = ($item[1] | split(" ")[0] | tonumber) * 1024 * 1024 * 1024 in . + {(.currentType): memoryBytes} | del(.currentType) else . end else . end )
使用说明
- 本地运行时,直接用
jq --stream -f script.jq input.json替代(. | tostream),内存占用会大幅降低。 - 该表达式会遍历流中的数据,仅在遇到
instanceType时暂存,遇到对应memory时计算字节数并添加到结果对象,完成后清除暂存的类型,避免内存堆积。
内容的提问来源于stack exchange,提问作者Jayson Reis
相关产品推荐
相关产品推荐

