从S3读取数据传入API遇ModelStateInvalid错误,硬编码数据正常
问题:从S3读取数据传入API触发验证错误
错误信息:
"error": {"code": "ModelStateInvalid", "message": "The request has exceeded the maximum number of validation errors.", "target": "HttpRequest"
正常工作的代码(直接构造字典)
def create_doc(self,client): self.n_docs = int(self.n_docs) document = {'addresses': {'SingleLocation': {'city': 'ABC', 'country': 'US', 'line1': 'Main', 'postalCode': '00000', 'region': 'CA' } }, 'commit': False, } response = client.cr_transc(document) jsn = response.json()
触发错误的代码(从S3读取数据)
def create_doc(self,client): self.n_docs = int(self.n_docs) document = data_from_s3() response = client.cr_transc(document) jsn = response.json() def data_from_s3(self): s3 = S3Hook() data = s3.read_key(bucket_name = self.bucket_name, key = self.data_key) return data
解决方法
核心原因
Airflow的S3Hook.read_key()返回的是字符串类型的文件内容,而API要求传入的是Python字典对象。直接传递字符串会导致API无法正确解析请求结构,触发批量验证错误。
修改后的代码
import json def create_doc(self,client): self.n_docs = int(self.n_docs) document = self.data_from_s3() # 补上self,修复类方法调用问题 response = client.cr_transc(document) jsn = response.json() def data_from_s3(self): s3 = S3Hook() data_str = s3.read_key(bucket_name = self.bucket_name, key = self.data_key) # 将JSON字符串解析为Python字典 return json.loads(data_str)
额外说明
原代码中create_doc方法调用data_from_s3()时缺少self前缀,这会导致无法正确调用类的成员方法,必须补上self.data_from_s3()才能正常执行。
内容的提问来源于stack exchange,提问作者Karthik
相关产品推荐
相关产品推荐

