递归遍历目录下所有TXT文件,提取并校验上下行速率数据
需求实现与代码修正
目录结构
- /home/python - file1.sh - file2.py - dirA - file1.txt - file2.txt - dirB - file3.txt - file5.txt - another - file6.txt - file10.txt - file11
TXT文件格式示例
所有TXT文件格式统一,以>开头的行分割为两部分:
> this is first output and some another contents these are some test lines to fill the file Testing download speed Download: 0.00 Mbit/s Testing upload speed Upload: 0.00 Mbit/s > this is second output but other texts go here too these are some test lines to fill the file Testing download speed Download: 1200.58 Mbit/s Testing upload speed Upload: 857.25 Kbit/s
处理逻辑
所有TXT文件均包含> this is first output和> this is second output两个标记段落,处理规则如下:
第一段落(对应Download1、Upload1)
- 定位
this is first output段落内的Download行:- 单位为Mbit时:
- 数值为0:输出
Download1 speed of {filename} is zero. - 数值小于600:输出
Download1 speed of {filename} less than 600 Mbits. - 数值≥600:输出
Download1 speed of {filename} is {value} Mbit/s.
- 数值为0:输出
- 单位为Kbit时:输出
Download1 speed of {filename} is {value} Kbit/s.
- 单位为Mbit时:
- 定位同段落内的
Upload行:- 单位为Mbit时:
- 数值为0:输出
Upload1 speed of {filename} is zero. - 数值小于600:输出
Upload1 speed of {filename} less than 600 Mbits. - 数值≥600:输出
Upload1 speed of {filename} is {value} Mbit/s.
- 数值为0:输出
- 单位为Kbit时:输出
Upload1 speed of {filename} is {value} Kbit/s.
- 单位为Mbit时:
第二段落(对应Download2、Upload2)
逻辑与第一段落完全一致,仅输出前缀替换为Download2、Upload2。
当前代码问题
现有代码无法区分两个段落的速度行,未实现完整的数值判断逻辑,且频繁重复打开日志文件会降低运行效率。
修正后的代码
#!/usr/bin/python3 from pathlib import Path def process_speed_line(line, prefix, filename, log_file): """处理单条Download/Upload行,生成对应输出内容""" parts = line.strip().split() value = float(parts[1]) unit = parts[2] if unit == 'Mbit/s': if value == 0.00: log_file.write(f"{prefix} speed of {filename} is zero.\n") elif value < 600: log_file.write(f"{prefix} speed of {filename} less than 600 Mbits.\n") else: log_file.write(f"{prefix} speed of {filename} is {value} Mbit/s.\n") elif unit == 'Kbit/s': log_file.write(f"{prefix} speed of {filename} is {value} Kbit/s.\n") # 遍历目标目录下所有txt文件 txt_files = list(Path("/home/python").rglob("*.txt")) # 仅打开一次日志文件,减少IO开销 with open('log.txt', 'w') as log: for file_path in txt_files: # 转换为相对路径,匹配期望输出格式 rel_path = file_path.relative_to("/home/python") current_section = 0 with open(file_path, 'r') as f: for line in f: line = line.strip() if line.startswith("> this is first output"): current_section = 1 elif line.startswith("> this is second output"): current_section = 2 elif current_section == 1 and line.startswith("Download:"): process_speed_line(line, "Download1", rel_path, log) elif current_section == 1 and line.startswith("Upload:"): process_speed_line(line, "Upload1", rel_path, log) elif current_section == 2 and line.startswith("Download:"): process_speed_line(line, "Download2", rel_path, log) elif current_section == 2 and line.startswith("Upload:"): process_speed_line(line, "Upload2", rel_path, log)
代码说明
- 段落区分:通过检测标记行确定当前处理的是第一还是第二段落,对应不同的输出前缀。
- 逻辑封装:将单条速度行的处理逻辑封装为函数,避免代码重复。
- 高效IO:日志文件仅打开一次,减少文件操作的性能损耗。
- 路径处理:使用
relative_to获取文件相对路径,匹配期望输出的格式。
期望输出示例
Download1 speed of dirA/file1.txt is zero. Upload1 speed of dirA/file1.txt is zero. Download2 speed of dirA/file1.txt is 1200.58 Mbit/s. Upload2 speed of dirA/file1.txt is 857.25 Kbit/s.
内容的提问来源于stack exchange,提问作者NomeEoleSem
相关产品推荐
相关产品推荐

