You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于时间更新多变量:DataFrame组件状态维护技术问询

组件在线状态维护:代码逻辑验证与优化

需求说明

我有个带Duration、EventLog、StartTime、FinishTime四列的DataFrame,要根据EventLog里的事件(组件故障/修复)更新对应组件的in_service状态——故障就设为False,修复就设为True,以此维护各组件的在线状态。

示例DataFrame:

DurationEventLogStartTimeFinishTime
12Component 1 repaired012
5Component 3 Failed05
12Component 5 Failed1224
44Component 1 Repaired5599

我写了一段代码实现,但不确定逻辑对不对,也想看看有没有优化空间:

import re

for i in range(8760):
    if 'Repaired' in data.loc[i,'EventLog'] :
        val = int(re.search(r'\d+', data.loc[i,'EventLog']).group())
        if val >= 0 and val <= 9:
            net.gen.loc[val,"in_service"] = True
        else:
            val = val - 10
            net.line.loc[val,"in_service"] = True        
        
    elif 'Failed' in data.loc[i,'EventLog'] :
        val = int(re.search(r'\d+', data.loc[i,'EventLog']).group())
        if val >= 0 and val <= 9:
            net.gen.loc[val,"in_service"] = False
        else:
            val = val - 10
            net.line.loc[val,"in_service"] = False

现有代码的问题点

  • 硬编码循环次数风险高:直接写range(8760)遍历,若DataFrame行数不是8760,要么行数不足触发索引越界错误,要么行数过多遗漏后续事件。
  • 正则匹配冗余:故障和修复分支重复执行re.search操作,浪费性能。
  • 大小写敏感漏判:示例中EventLog存在repaired(小写)和Repaired(大写)两种格式,现有代码用'Repaired'判断会漏掉小写事件,导致状态不更新。
  • 索引假设太绝对:用loc[val]默认组件编号与DataFrame行索引完全对应,若组件编号不连续或索引非整数,直接触发错误。
  • 无异常处理:若EventLog无数字或格式异常,re.search返回None,调用.group()会直接抛出AttributeError。

优化后的实现方案

方案1:安全遍历+异常容错

直接遍历DataFrame行,兼容各种边界情况:

import re

for idx, row in data.iterrows():
    event = row['EventLog']
    # 统一转小写判断,兼容大小写差异
    is_repaired = 'repaired' in event.lower()
    is_failed = 'failed' in event.lower()
    
    # 非目标事件直接跳过
    if not (is_repaired or is_failed):
        continue
    
    # 提取组件编号,增加容错逻辑
    match = re.search(r'\d+', event)
    if not match:
        print(f"事件「{event}」无法提取组件编号,跳过")
        continue
    comp_num = int(match.group())
    
    # 确定要设置的状态值
    status = True if is_repaired else False
    
    # 选择对应数据集
    target_df = net.gen if 0 <= comp_num <=9 else net.line
    if comp_num > 9:
        comp_num -= 10
    
    # 验证组件编号存在后再更新,避免索引错误
    if comp_num in target_df.index:
        target_df.loc[comp_num, 'in_service'] = status
    else:
        print(f"组件编号{comp_num}在对应数据集中不存在,跳过")

方案2:预处理事件+批量处理(高效版)

先解析所有事件再批量更新,适合数据量大的场景:

import re
import pandas as pd

# 定义事件解析函数
def parse_event(event_str):
    event_lower = event_str.lower()
    num_match = re.search(r'\d+', event_str)
    if not num_match:
        return None, None
    comp_num = int(num_match.group())
    if 'repaired' in event_lower:
        return comp_num, 'repaired'
    elif 'failed' in event_lower:
        return comp_num, 'failed'
    else:
        return None, None

# 为DataFrame新增解析结果列
data[['comp_num', 'event_type']] = data['EventLog'].apply(lambda x: pd.Series(parse_event(x)))

# 过滤无效事件并按时间排序(确保状态更新符合时间逻辑)
valid_events = data.dropna(subset=['comp_num', 'event_type']).sort_values('StartTime')

# 批量更新组件状态
for _, row in valid_events.iterrows():
    comp_num = row['comp_num']
    status = True if row['event_type'] == 'repaired' else False
    
    target_df = net.gen if 0 <= comp_num <=9 else net.line
    if comp_num >9:
        comp_num -=10
    
    if comp_num in target_df.index:
        target_df.loc[comp_num, 'in_service'] = status

额外注意事项

  • 务必先将DataFrame按StartTime排序,否则事件乱序会导致状态更新逻辑错误。
  • 提前确认net.gen和net.line的索引包含所有可能的组件编号,减少“组件不存在”的报错。
  • 可将print替换为日志模块(如logging),方便后续问题排查。

内容的提问来源于stack exchange,提问作者Giorgos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 07:55:30