基于时间更新多变量:DataFrame组件状态维护技术问询
组件在线状态维护:代码逻辑验证与优化
需求说明
我有个带Duration、EventLog、StartTime、FinishTime四列的DataFrame,要根据EventLog里的事件(组件故障/修复)更新对应组件的in_service状态——故障就设为False,修复就设为True,以此维护各组件的在线状态。
示例DataFrame:
| Duration | EventLog | StartTime | FinishTime |
|---|---|---|---|
| 12 | Component 1 repaired | 0 | 12 |
| 5 | Component 3 Failed | 0 | 5 |
| 12 | Component 5 Failed | 12 | 24 |
| 44 | Component 1 Repaired | 55 | 99 |
我写了一段代码实现,但不确定逻辑对不对,也想看看有没有优化空间:
import re for i in range(8760): if 'Repaired' in data.loc[i,'EventLog'] : val = int(re.search(r'\d+', data.loc[i,'EventLog']).group()) if val >= 0 and val <= 9: net.gen.loc[val,"in_service"] = True else: val = val - 10 net.line.loc[val,"in_service"] = True elif 'Failed' in data.loc[i,'EventLog'] : val = int(re.search(r'\d+', data.loc[i,'EventLog']).group()) if val >= 0 and val <= 9: net.gen.loc[val,"in_service"] = False else: val = val - 10 net.line.loc[val,"in_service"] = False
现有代码的问题点
- 硬编码循环次数风险高:直接写
range(8760)遍历,若DataFrame行数不是8760,要么行数不足触发索引越界错误,要么行数过多遗漏后续事件。 - 正则匹配冗余:故障和修复分支重复执行
re.search操作,浪费性能。 - 大小写敏感漏判:示例中
EventLog存在repaired(小写)和Repaired(大写)两种格式,现有代码用'Repaired'判断会漏掉小写事件,导致状态不更新。 - 索引假设太绝对:用
loc[val]默认组件编号与DataFrame行索引完全对应,若组件编号不连续或索引非整数,直接触发错误。 - 无异常处理:若
EventLog无数字或格式异常,re.search返回None,调用.group()会直接抛出AttributeError。
优化后的实现方案
方案1:安全遍历+异常容错
直接遍历DataFrame行,兼容各种边界情况:
import re for idx, row in data.iterrows(): event = row['EventLog'] # 统一转小写判断,兼容大小写差异 is_repaired = 'repaired' in event.lower() is_failed = 'failed' in event.lower() # 非目标事件直接跳过 if not (is_repaired or is_failed): continue # 提取组件编号,增加容错逻辑 match = re.search(r'\d+', event) if not match: print(f"事件「{event}」无法提取组件编号,跳过") continue comp_num = int(match.group()) # 确定要设置的状态值 status = True if is_repaired else False # 选择对应数据集 target_df = net.gen if 0 <= comp_num <=9 else net.line if comp_num > 9: comp_num -= 10 # 验证组件编号存在后再更新,避免索引错误 if comp_num in target_df.index: target_df.loc[comp_num, 'in_service'] = status else: print(f"组件编号{comp_num}在对应数据集中不存在,跳过")
方案2:预处理事件+批量处理(高效版)
先解析所有事件再批量更新,适合数据量大的场景:
import re import pandas as pd # 定义事件解析函数 def parse_event(event_str): event_lower = event_str.lower() num_match = re.search(r'\d+', event_str) if not num_match: return None, None comp_num = int(num_match.group()) if 'repaired' in event_lower: return comp_num, 'repaired' elif 'failed' in event_lower: return comp_num, 'failed' else: return None, None # 为DataFrame新增解析结果列 data[['comp_num', 'event_type']] = data['EventLog'].apply(lambda x: pd.Series(parse_event(x))) # 过滤无效事件并按时间排序(确保状态更新符合时间逻辑) valid_events = data.dropna(subset=['comp_num', 'event_type']).sort_values('StartTime') # 批量更新组件状态 for _, row in valid_events.iterrows(): comp_num = row['comp_num'] status = True if row['event_type'] == 'repaired' else False target_df = net.gen if 0 <= comp_num <=9 else net.line if comp_num >9: comp_num -=10 if comp_num in target_df.index: target_df.loc[comp_num, 'in_service'] = status
额外注意事项
- 务必先将DataFrame按
StartTime排序,否则事件乱序会导致状态更新逻辑错误。 - 提前确认
net.gen和net.line的索引包含所有可能的组件编号,减少“组件不存在”的报错。 - 可将
print替换为日志模块(如logging),方便后续问题排查。
内容的提问来源于stack exchange,提问作者Giorgos
相关产品推荐
相关产品推荐

