EKS集群部署Loki+S3存储遇IRSA凭证获取405错误求助
在EKS集群中部署Loki(使用grafana/loki 6.19.0 Chart,Loki版本3.2.1),采用S3桶作为存储介质,已配置IAM Roles for Service Accounts(IRSA)授权Loki访问S3,但Loki日志出现以下错误:
WebIdentityErr: failed to retrieve credentials caused by: SerializationError: failed to unmarshal error message status code: 405
完整错误详情:
level=error ts=2024-11-14T10:08:39.203006213Z caller=flush.go:261
component=ingester loop=1 org_id=1 msg="failed to flush" retries=2
err="failed to flush chunks: store put chunk: WebIdentityErr: failed
to retrieve credentials
caused by: SerializationError: failed to
unmarshal error message
status code: 405, request id:
caused by:
UnmarshalError: failed to unmarshal error message
00000000 3c 3f
78 6d 6c 20 76 65 72 73 69 6f 6e 3d 22 31 |.<C|
00000030 6f 64 65 3e 0d 65 74 68
6f 64 4e 6f 74 41 6c 6c |ode>MethodNotAll|
00000040 6f 77 65 64 3c
2f 43 6f 64 65 3e 3c 4d 65 73 73 |owed<Mess|
00000050 61 67
age>The 00000060 64 20 6d 65 74 68 6f 64 20 69 73 20 6e 6f 74 20
|d method is not |
00000070 61 6c 6c 6f 77 65 64 20 61 67 61 69 6e
73 74 20 |allowed against |
00000080 74 68 69 73 20 72 65 73 6f 75
72 63 65 2e 3c 2f |this resource.</|
00000090 4d 65 73 73 61 67 65
3e 3c 4d 65 74 68 6f 64 3e |Message>|
000000a0 50 4f 53 54
3c 2f 4d 65 74 68 6f 64 3e 3c 52 65 |POST<Re|
000000b0 73
6f 75 72 63 65 54 79 70 65 3e 53 45 52 56 49
|sourceType>SERVI|
000000c0 43 45 3c 2f 52 65 73 6f 75 72 63 65 54
79 70 65 |CE</ResourceType|
000000d0 3e 3c 52 65 71 75 65 73 74 49
64 3e 45 36 50 34 |>E6P4|
000000e0 34 34 34 43 51 32 4e
30 35 45 52 4a 3c 2f 52 65 |444CQ2N05ERJ</Re|
000000f0 71 75 65 73
74 49 64 3e 3c 48 6f 73 74 49 64 3e |questId>|
00000100 70
4f 30 35 58 6d 42 4b 33 6d 45 36 30 43 32 66
|pO05XmBK3mE60C2f|
00000110 6c 56 74 36 76 65 45 36 37 41 6c 56 43
58 4a 4d |lVt6veE67AlVCXJM|
00000120 61 9c 66 39 65 50 36 67 56 53
7a 70 57 49 78 50 |aLf9eP6gVSzpWIxP|
00000130 57 32 32 63 58 76 51
51 73 6b 54 6b 35 78 64 6d |W22cXvQQskTk5xdm|
00000140 62 52 6b 36
35 73 48 33 4e 39 2f 54 6e 39 77 30 |bRk65sH3N9/Tn9w0|
00000150 46
72 31 2f 2f 75 61 55 34 6d 2b 78 64 6f 46 59
|Fr1//uaU4m+xdoFY|
00000160 71 7a 6e 47 79 86 35 51 4a 2b 59 3d 3c
2f 48 6f |qznGyF5QJ+Y=</Ho|
00000170 73 74 49 64 3e 3c 2f 45 72 72
6f 72 3e |stId>|caused by: unknown error
response tag, {{ Error} []}, num_chunks: 1, labels:
{app="curl-test", container="curl-test",
filename="/var/log/pods/default_curl-test_5cb0fb1b-922b-4a1c-b8e0-e6506bec3cee/curl-test/239.log",
job="default/curl-test", namespace="default",
node_name="ip-172-32-18-62.ap-southeast-1.compute.internal",
pod="curl-test", service_name="curl-test", stream="stdout"}"
当前使用的loki-values.yml配置如下:
loki: auth-enabled: false image: tag: 3.2.1 schemaConfig: configs: - from: "2024-04-01" store: tsdb object_store: s3 schema: v13 index: prefix: loki_index_ period: 24h annotations: eks.amazonaws.com/role-arn: arn:aws:iam::469030223850:role/loki-test-role-again storage_config: aws: # s3: s3://ap-southeast-1/chunks-buck region: ap-southeast-1 bucketnames: chunks-buck s3forcepathstyle: false pattern_ingester: enabled: true limits_config: allow_structured_metadata: true volume_enabled: true retention_period: 672h # 28 days retention querier: max_concurrent: 4 storage: bucketNames: chunks: chunks-buck ruler: ruler-buck admin: admin-buck type: s3 s3: s3: s3://ap-southeast-1/chunks-buck endpoint: s3.ap-southeast-1.amazonaws.com region: ap-southeast-1 s3ForcePathStyle: false memcached: chunk_cache: resources: requests: cpu: 500m memory: 256Mi # Lower memory request limits: memory: 512Mi # Lower memory limit deploymentMode: SimpleScalable serviceAccount: # -- Specifies whether a ServiceAccount should be created create: true # -- The name of the ServiceAccount to use. # If not set and create is true, a name is generated using the fullname template name: loki-sa # -- Image pull secrets for the service account imagePullSecrets: [] # -- Labels for the service account labels: {} # -- Annotations for the service account annotations: eks.amazonaws.com/role-arn: arn:aws:iam::469030223850:role/loki-test-role-again # -- Set this toggle to false to opt out of automounting API credentials for the service account automountServiceAccountToken: true backend: replicas: 3 persistence: volumeClaimsEnabled: false read: replicas: 3 persistence: storageClass: gp2 write: replicas: 3 persistence: volumeClaimsEnabled: false # -- Configuration for the read pod(s) minio: enabled: false singleBinary: persistence: storageClass: gp2 ingestor: resources: requests: memory: 512Mi limits: memory: 1Gi persistence: claims: - name: data storageClass: gp2 querier: persistence: storageClass: gp2 indexGateway: persistence: storageClass: gp2 compactor: persistence: storageClass: gp2 claims: - name: data storageClass: gp2 bloomGateway: persistence: claims: - name: data storageClass: gp2 bloomPlanner: persistence: claims: - name: data storageClass: gp2 patternIngester: persistence: storageClass: gp2 claims: - name: data storageClass: gp2 ruler: persistence: storageClass: gp2 resultsCache: persistence: storageClass: gp2 chunksCache: allocatedMemory: 1000 persistence: storageClass: gp2
从错误信息中的XML响应可以明确:Loki在尝试获取IRSA凭证时使用了POST方法,但该方法不被允许(MethodNotAllowed),这是核心问题。结合配置来看,有两个关键诱因:
- 重复的IRSA注解配置:在
loki.annotations和serviceAccount.annotations两处都配置了eks.amazonaws.com/role-arn,导致Pod注解冲突,干扰IRSA凭证获取流程。 - S3存储配置冗余冲突:同时配置了
loki.storage_config.aws和loki.storage.s3,这两个配置块针对Loki不同存储API版本,冗余配置会导致Loki使用错误的凭证请求路径。
1. 移除重复的IRSA注解
删除loki.annotations下的eks.amazonaws.com/role-arn配置,只保留serviceAccount.annotations中的该注解,确保IRSA配置唯一且正确关联到ServiceAccount。
修改后的loki配置片段:
loki: auth-enabled: false image: tag: 3.2.1 schemaConfig: configs: - from: "2024-04-01" store: tsdb object_store: s3 schema: v13 index: prefix: loki_index_ period: 24h # 移除此处重复的IRSA注解 # annotations: # eks.amazonaws.com/role-arn: arn:aws:iam::469030223850:role/loki-test-role-again storage_config: aws: # s3: s3://ap-southeast-1/chunks-buck region: ap-southeast-1 bucketnames: chunks-buck s3forcepathstyle: false # 其他配置保持不变
2. 统一S3存储配置
Loki 3.x版本支持新的统一存储API,建议保留loki.storage配置,删除旧的loki.storage_config块,避免配置冲突。同时删除loki.storage.s3中的s3和endpoint字段,AWS原生S3会自动根据region推导正确的endpoint。
修改后的loki存储配置片段:
loki: # 移除旧的storage_config配置块 # storage_config: # aws: # # s3: s3://ap-southeast-1/chunks-buck # region: ap-southeast-1 # bucketnames: chunks-buck # s3forcepathstyle: false storage: bucketNames: chunks: chunks-buck ruler: ruler-buck admin: admin-buck type: s3 s3: region: ap-southeast-1 s3ForcePathStyle: false # 其他配置保持不变
3. 验证IAM角色权限
确保loki-test-role-again角色具备以下S3操作权限:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:PutObject", "s3:GetObject", "s3:ListBucket", "s3:DeleteObject" ], "Resource": [ "arn:aws:s3:::chunks-buck", "arn:aws:s3:::chunks-buck/*", "arn:aws:s3:::ruler-buck", "arn:aws:s3:::ruler-buck/*", "arn:aws:s3:::admin-buck", "arn:aws:s3:::admin-buck/*" ] } ] }
4. 重新部署Loki
执行Helm升级命令重新部署Loki:
helm upgrade loki grafana/loki -f loki-values.yml -n <你的命名空间>
内容的提问来源于stack exchange,提问作者Saksham Paliwal

