如何在Python中从API返回的大文本中提取指定子字符串?
问题场景
API返回的任务执行文本格式如下:
TASK [Do this] OK: { "changed":false, "msg": "check ok" } TASK [Do that] OK TASK [Do x] Fatal: "Error message x" TASK [Do y] OK TASK [Do z] Fatal: "Stopped because of previous error"
需提取Error message x这类实际错误信息,现有请求代码:
url = # API URL r = requests.get(url, verify=False, allow_redirects=True, headers=headers, timeout=10) output = r.text
使用output.split("Fatal", 1)[1]时触发list index out of range错误,且文本存在大量转义换行格式问题。
解决方案
1. 清理文本格式
API返回的文本可能包含转义换行符(如\n),先替换为实际换行:
cleaned_output = output.replace('\\n', '\n')
2. 正则精准提取目标错误
通过正则匹配所有Fatal:后的内容,再过滤掉连锁错误(如Stopped because of previous error):
import re # 匹配Fatal: "包裹的错误内容 pattern = r'Fatal: "([^"]+)"' all_errors = re.findall(pattern, cleaned_output) # 筛选出非连锁的目标错误 target_errors = [err for err in all_errors if not err.startswith('Stopped because of previous error')]
3. 避免索引越界错误
直接使用split会在无Fatal内容时抛出索引错误,先判断存在性再处理:
if not target_errors: print("未找到目标错误信息") else: for err in target_errors: print(f"- {err}")
完整代码示例
import requests import re url = # 你的API URL headers = {} # 你的请求头配置 r = requests.get(url, verify=False, allow_redirects=True, headers=headers, timeout=10) output = r.text # 清理转义换行 cleaned_output = output.replace('\\n', '\n') # 提取并过滤错误 pattern = r'Fatal: "([^"]+)"' all_errors = re.findall(pattern, cleaned_output) target_errors = [err for err in all_errors if not err.startswith('Stopped because of previous error')] # 输出结果 if target_errors: print("提取到的错误信息:") for err in target_errors: print(f"- {err}") else: print("未找到目标错误信息")
内容的提问来源于stack exchange,提问作者Zodi
相关产品推荐
相关产品推荐

