You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让subprocess.check_output丢弃旧内容仅保留最后N字节?

解决子进程超大输出仅保留最后N字节的高效方案

核心问题

使用subprocess.check_output处理每秒输出2GB+的命令时,PIPE无大小限制会耗尽内存,仅需保留最后N字节,现有循环读取方案速度跟不上输出速率且浪费内存,需要能快速丢弃旧内容的IO机制。

最优方案:利用系统tail命令(推荐)

借助Linux系统自带的tail命令,由内核层面直接截断数据流,只保留最后N字节,完全避免用户空间内存堆积,速度能匹配高速输出:

import subprocess

N = 5
# 替换为你的目标命令
target_cmd = "dd if=/dev/zero bs=10M count=1000 && echo test"

# 通过shell管道让tail过滤最后N字节
with subprocess.Popen(
    ["sh", "-c", f"{target_cmd} | tail -c {N}"],
    stdout=subprocess.PIPE,
    stderr=subprocess.DEVNULL  # 若需要 stderr 可移除该行或改为 subprocess.PIPE
) as proc:
    last_bytes = proc.stdout.read()
    print(last_bytes)

优势

  • 速度极快:tail是系统级优化工具,直接在内核管道中处理数据,无需将全部输出读到用户空间
  • 内存占用极低:仅保留最后N字节,不会堆积大量数据
  • 实现简单:无需手动处理循环读取和数据截断逻辑

纯Python实现方案(跨平台/无外部命令需求)

如果无法依赖外部命令,可采用非阻塞读取+环形缓冲区的方式,主动丢弃超出N字节的旧数据,平衡速度与内存占用:

import subprocess
import os
import fcntl
import time

def set_nonblocking(fd):
    """将文件描述符设置为非阻塞模式"""
    flags = fcntl.fcntl(fd, fcntl.F_GETFL)
    fcntl.fcntl(fd, fcntl.F_SETFL, flags | os.O_NONBLOCK)

N = 5
target_cmd = "dd if=/dev/zero bs=10M count=1000 && echo test"
block_size = 1024 * 1024  # 1MB读取块,可根据输出速率调整
ring_buffer = b""

with subprocess.Popen(target_cmd, shell=True, stdout=subprocess.PIPE) as proc:
    set_nonblocking(proc.stdout.fileno())
    
    # 进程运行时循环读取数据
    while proc.poll() is None:
        try:
            chunk = proc.stdout.read(block_size)
            if chunk:
                # 只保留最后N字节,自动丢弃前面的数据
                ring_buffer = (ring_buffer + chunk)[-N:]
        except BlockingIOError:
            # 无数据时短暂休眠,避免CPU空转
            time.sleep(0.001)
    
    # 读取进程结束后剩余的最后一批数据
    while True:
        try:
            remaining = proc.stdout.read(block_size)
            if not remaining:
                break
            ring_buffer = (ring_buffer + remaining)[-N:]
        except BlockingIOError:
            break

print(ring_buffer)

注意事项

  • 调整block_size:块越大,系统调用次数越少,速度越快;块过小会增加CPU开销
  • 非阻塞模式:避免读取操作阻塞主线程,保证能跟上高速输出
  • 环形缓冲区:始终只保留最后N字节,前面的数据直接被覆盖丢弃,内存占用固定为N字节

不推荐的方案

  • 临时文件存储再读取:磁盘IO速度远低于内存管道,无法匹配2GB/s的输出速率,会导致进程阻塞或磁盘空间占用过高
  • 尝试seek定位stdout:stdout是管道流,不支持随机定位,此方案完全不可行

内容的提问来源于stack exchange,提问作者Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 06:10:52