如何让subprocess.Popen.communicate的输出按行保留?
问题:subprocess输出按行处理并写入DataFrame失效的解决方法
将subprocess.Popen.communicate的STDOUT写入pd.DataFrame时,简单输出处理正常,但添加行过滤逻辑(比如line.startswith('This'))后失效——原因是直接遍历解码后的grep_stdout字符串时,会逐个字符迭代而非按行分割。
解决思路
核心是把解码后的字符串按行拆分,确保遍历的是完整的行而非单个字符,以下是两种常用方法:
方法1:用splitlines()分割成行列表
通过splitlines()将字符串按换行符分割成独立行的列表,遍历每行时添加过滤逻辑,注意写入时补回换行符(splitlines()会移除原换行符)。
修改后的代码:
import subprocess import io import pandas as pd strings = ['Hello\tWorld!', 'This\tis', 'a\tTest!'] string = '\n'.join(strings) cmd_grep = ['grep', 's'] process_grep = subprocess.Popen(cmd_grep, stdin=subprocess.PIPE, stdout=subprocess.PIPE) grep_stdout = process_grep.communicate(input=string.encode('utf-8'))[0].decode('utf-8') grep_csv = io.StringIO() # 按行分割字符串,遍历每一行 for line in grep_stdout.splitlines(): # 添加行过滤逻辑 if line.startswith('This'): grep_csv.write(line + '\n') # 补回换行符,保证DataFrame读取格式正确 grep_csv.seek(0) grep_results = pd.read_csv(grep_csv, sep='\t', header=None, names=['Word1', 'Word2']) grep_csv.close() print(grep_results)
方法2:用io.StringIO直接包裹输出后按行读取
把解码后的字符串直接传入io.StringIO,此时迭代该对象会自动按行返回内容,无需手动处理换行符。
修改后的代码:
import subprocess import io import pandas as pd strings = ['Hello\tWorld!', 'This\tis', 'a\tTest!'] string = '\n'.join(strings) cmd_grep = ['grep', 's'] process_grep = subprocess.Popen(cmd_grep, stdin=subprocess.PIPE, stdout=subprocess.PIPE) grep_stdout = process_grep.communicate(input=string.encode('utf-8'))[0].decode('utf-8') with io.StringIO(grep_stdout) as grep_io: grep_csv = io.StringIO() # 直接遍历StringIO对象,每次迭代返回一行 for line in grep_io: if line.startswith('This'): grep_csv.write(line) grep_csv.seek(0) grep_results = pd.read_csv(grep_csv, sep='\t', header=None, names=['Word1', 'Word2']) print(grep_results)
内容的提问来源于stack exchange,提问作者gernophil
相关产品推荐
相关产品推荐

