Python提取HTTP响应JSON指定字段写入本地文件方案
问题场景
- 调用查询接口:
https://test.url/api/v2/services/level/app?domainname={}&offset=0&limit=100,从domains.txt文件逐行读取待查询域名发起请求 - 初始查询实现代码:
with open('domains.txt') as f_input: for id in f_input: url = "https://test.url/api/v2/services/level/app?domainname={}&offset=0&limit=100".format(id.strip()) resp = requests.get(url, headers={'Accept':'application/json','Api-Token': 'XXXX'}) json_string = json.dumps(resp.json(), sort_keys=True, indent=4, separators=(',', ': ')) print(json_string)
domains.txt文件内容:
url1.com url2.com url3.com
- 接口返回JSON结构(补全外层缺失的大括号):
{ "data": [ { "app_name": "App1", "category_name": "Category1", "level": 44, "indicator": "poor", "id": 9563 } ], "status": "Success", "status_code": 200, "total_query_count": 1 }
- 需求:将查询对应的域名、响应中的
app_name、category_name字段写入ergebnis.txt,每行格式为域名 App_Name Category。
遇到的报错
最初编写的写入代码运行抛出异常:TypeError: 'types.GenericAlias' object is not iterable,原错误代码如下:
with open("ergebnis.txt", "w") as f: for i in list['data']: f.write("{0} {1}\n".format(i["app_name"],str(i["category_name"])))
报错原因:代码中直接对Python内置类型list取下标['data'],没有引用实际存储接口返回数据的响应对象,内置类型本身不可迭代,因此抛出类型错误。
后续调整代码后,已经可以将响应解析为resp_json对象,遍历resp_json["data"]字段、以追加模式写入app_name和category_name两个字段,但缺少对应关联的查询域名,需要补全三个字段的拼接写入。
修复后完整代码
import requests import json # 提前以写入模式打开结果文件,避免循环内反复操作文件IO with open("ergebnis.txt", "w", encoding="utf-8") as f_out: with open('domains.txt', 'r', encoding='utf-8') as f_in: for line in f_in: # 去除每行首尾换行、空白字符,得到当前查询的域名 current_domain = line.strip() # 跳过空行 if not current_domain: continue # 拼接请求地址 url = f"https://test.url/api/v2/services/level/app?domainname={current_domain}&offset=0&limit=100" # 发起GET请求 resp = requests.get( url, headers={ 'Accept': 'application/json', 'Api-Token': 'XXXX' # 替换为实际可用的Api Token } ) # 解析响应为JSON对象 resp_json = resp.json() # 遍历data下的所有条目 for item in resp_json.get("data", []): app_name = item.get("app_name", "") category_name = item.get("category_name", "") # 按要求拼接三个字段写入文件,空格分隔 f_out.write(f"{current_domain} {app_name} {category_name}\n")
关键修改说明
- 遍历域名列表时,用
current_domain变量存储当前正在查询的域名,写文件时直接引用该变量即可关联对应域名,不需要额外做映射 - 废弃错误的
list['data']写法,直接对解析后的响应对象resp_json取data字段遍历,同时用get方法加空列表默认值,避免接口返回异常时抛出KeyError - 统一在循环外打开结果文件使用覆盖写入模式,无需在循环内反复以追加模式开关文件,执行效率更高
- 增加空行跳过、字段默认值兜底逻辑,避免文件空行、接口返回字段缺失导致的运行中断
内容的提问来源于stack exchange,提问作者Jack
相关产品推荐
相关产品推荐

