You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python requests库获取大日志时response.text卡顿求助

问题描述

使用Python的requests库获取大体积日志并提取特定字符串时,大部分日志能正常打印内容、文本及长度,但有一份日志卡在return response.text步骤,无法继续执行。更大体积的日志反而能正常处理,不确定原因,寻求分析与解决思路。

代码示例

def fetch_webpage_content(url):
    try:
        response = requests.get(url, headers = my_headers)
        print("got it")
        print(len(response.content))
        #print(response.text)
        #print(response.content)
        response.raise_for_status()
        return response.text
    except requests.RequestException as e:
        print(f"An error occurred while fetching web page content: {e}")
        return None

def process_webpage(url, pattern):
    text = fetch_webpage_content(url)
    print("in process_webpage")
    if text is not None:
        match = search_pattern(pattern, text)
        if match:
            return match.group()
        return ''
    else:
        return ''

输出对比

正常输出(输出1)

got it
33131960
in process_webpage
got it
9383326
in process_webpage

got it
33131960
in process_webpage
got it
13795885
in process_webpage

异常输出(输出2)

got it
33131960
in process_webpage
got it
25565370
分析与解决思路
  • 排查编码问题:response.text会自动根据响应头猜测编码,若目标日志的编码不标准(比如非UTF-8且无正确标识),requests可能在解码时卡顿。可以手动指定编码(如response.encoding = 'utf-8')后再获取text,或者直接用response.content拿到字节流,自行处理解码逻辑。
  • 改用流式处理:没必要一次性加载全量日志转成字符串,开启requests的流式请求(stream=True),分块读取内容并匹配正则,既节省内存,也能避免全量解码的卡顿。示例代码:
def fetch_webpage_content_stream(url, pattern):
    try:
        with requests.get(url, headers=my_headers, stream=True) as response:
            response.raise_for_status()
            buffer = b''
            chunk_size = 8192
            for chunk in response.iter_content(chunk_size=chunk_size):
                buffer += chunk
                # 尝试在当前缓冲区匹配目标模式
                match = search_pattern(pattern, buffer.decode('utf-8', errors='ignore'))
                if match:
                    return match.group()
                # 缓冲区过大时截断,仅保留可能匹配的尾部内容
                if len(buffer) > len(pattern) * 2:
                    buffer = buffer[-len(pattern)*2:]
        return ''
    except requests.RequestException as e:
        print(f"Error fetching content: {e}")
        return None
  • 检查响应内容特殊性:那份异常日志可能包含特殊字符、未闭合的字节序列,或是二进制格式伪装成文本。可以打印response.content的前几百字节,查看是否有非文本标记、乱码片段,调整解码策略。
  • 优化超时与异常捕获:当前代码未设置超时,若服务器传输最后部分数据延迟,可能导致卡顿。给requests.get添加timeout参数(如timeout=30),同时单独捕获解码异常:
try:
    return response.text
except UnicodeDecodeError as e:
    print(f"Decode error: {e}")
    # 退而求其次返回替换乱码后的内容
    return response.content.decode('utf-8', errors='replace')
  • 监控内存占用:不同日志的内容结构可能导致内存占用差异,用memory_profiler等工具监控处理异常日志时的内存变化,排查是否因内存飙升导致系统卡顿。

内容的提问来源于stack exchange,提问作者leaTheBest

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 02:45:28