Python列表循环处理:文件读取、注释行移除及时间转换问题
处理in.dat文件:移除注释、清理换行符及时间转换
需求与问题
需要完成以下操作:
- 读取
in.dat文件 - 移除所有以
#开头的注释行 - 将剩余行处理为元组并存入列表
- 将每行的开始/结束时间转换为距午夜的时长(分钟)
当前遇到的问题:
- 注释行移除功能逻辑错误,无法过滤注释
- 不知道如何处理行尾的
\n换行符 - 不清楚如何将12小时制时间转换为距午夜的分钟数
in.dat文件内容
# read from file data about one day # format: start_time:end_time:#steps 09.30AM:09.45AM:220 11.45AM:12.23PM:300 11.45AM:10.23AM:302 2.45PM:3.23PM:202 3.45PM:3.53PM:90 5.45PM:5.53PM:80 6.45PM:7.23PM:1000 10.45PM:10.53PM:102
当前代码(存在问题)
# Reads file and returns a list of lines in string format def read_data(fname): with open("in.dat", "r") as f: data = f.readlines() f.close() # 多余:with语句会自动关闭文件 return data # Takes a list of lines as input and returns a new list of lines with the # comment lines (the ones that begin with #) removed. def remove_comment_lines(data): result = [] for name in data: if len(name) <=30: # 判断逻辑错误:应该检查是否以#开头 result.append(data) # 错误:应该添加当前行name而非整个data列表 print(result) return result # 错误:循环第一次就返回,未遍历所有行
调用语句:
remove_comment_lines(read_data("in.dat"))
解决方案
1. 修复注释移除与换行符处理
首先修正read_data和注释过滤逻辑,同时处理换行符并将行分割为元组:
def read_data(fname): # with语句自动管理文件关闭,无需手动调用f.close() with open(fname, "r") as f: return f.readlines() def process_raw_data(raw_data): processed = [] for line in raw_data: # 先移除行尾的换行符及首尾空白字符 stripped_line = line.strip() # 跳过空行和以#开头的注释行 if not stripped_line or stripped_line.startswith("#"): continue # 按冒号分割为元组,并存入列表 parts = tuple(stripped_line.split(":")) processed.append(parts) return processed
2. 时间转换:12小时制转距午夜的分钟数
写一个工具函数,将09.30AM、2.45PM这类格式的时间转换为距午夜的总分钟数:
def time_to_minutes(time_str): # 拆分时间与AM/PM标识 period = time_str[-2:] time_part = time_str[:-2] # 拆分小时和分钟 hour_str, minute_str = time_part.split(".") hour = int(hour_str) minute = int(minute_str) # 12小时制转24小时制 if period == "PM": if hour != 12: hour += 12 else: # AM if hour == 12: hour = 0 # 计算距午夜的总分钟数 return hour * 60 + minute
3. 完整流程代码
整合所有功能,最终得到包含时间转换的处理结果:
def read_data(fname): with open(fname, "r") as f: return f.readlines() def process_raw_data(raw_data): processed = [] for line in raw_data: stripped_line = line.strip() if not stripped_line or stripped_line.startswith("#"): continue start_time, end_time, steps = stripped_line.split(":") # 转换时间为分钟数,步数转为整数 start_min = time_to_minutes(start_time) end_min = time_to_minutes(end_time) processed.append((start_min, end_min, int(steps))) return processed def time_to_minutes(time_str): period = time_str[-2:] time_part = time_str[:-2] hour_str, minute_str = time_part.split(".") hour = int(hour_str) minute = int(minute_str) if period == "PM": if hour != 12: hour += 12 else: if hour == 12: hour = 0 return hour * 60 + minute # 调用示例 if __name__ == "__main__": raw_lines = read_data("in.dat") result = process_raw_data(raw_lines) print("处理后结果:") for item in result: print(item)
运行后输出示例:
处理后结果: (570, 585, 220) (705, 743, 300) (705, 623, 302) (885, 923, 202) (945, 953, 90) (1065, 1073, 80) (1185, 1223, 1000) (1365, 1373, 102)
内容的提问来源于stack exchange,提问作者user159
相关产品推荐
相关产品推荐

