You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析含十六进制编码的CAN日志文件并在Python中用pandas处理

CAN总线日志解析实现思路

你要处理的这类CAN总线日志结构非常规整,完全可以通过简单的正则匹配+字段拆分完成解析,最终输出pandas可直接处理的DataFrame格式,具体实现步骤如下:


1 明确日志结构拆分规则

整份日志可以拆为三类可解析的内容:

  • 头部元数据:以;$开头,存储文件版本、采集开始时间两个全局字段
  • 总线配置表:注释块中以数字开头的行,存储每个总线的编号、名称、连接地址、协议、波特率
  • 报文原始数据:以数字+)开头的行,存储每一条CAN报文的全量字段

2 核心解析代码实现

直接复制使用即可,注释已标注所有关键逻辑,只需要替换成你的本地日志文件路径:

import re
import pandas as pd

# 初始化存储容器
metadata = {}
bus_config = []
can_messages = []

# 预定义正则匹配规则,适配日志固定格式
meta_pattern = re.compile(r';\$(\w+)=([\d.]+)')
bus_pattern = re.compile(r';\s+(\d)\s+(\w+)\s+([\w@_]+)\s+(\w+)\s+(\d+\s+kbit/s)')
msg_pattern = re.compile(r'\s+(\d+)\)\s+([\d.]+)\s+(\d)\s+(\w+)\s+([0-9A-Fa-f]+)\s+-\s+(\d)\s+([0-9A-Fa-f\s]+)')

# 逐行读取并解析日志
with open('替换为你的CAN日志文件路径.log', 'r', encoding='utf-8') as f:
    for line in f:
        line = line.rstrip('\n')
        # 跳过空行和纯分隔线
        if not line.strip() or line.strip().startswith(';---'):
            continue
        
        # 匹配头部元数据
        meta_match = meta_pattern.match(line)
        if meta_match:
            key, val = meta_match.groups()
            metadata[key] = float(val) if '.' in val else int(val)
            continue
        
        # 匹配总线配置信息
        bus_match = bus_pattern.match(line)
        if bus_match:
            bus_no, name, conn, proto, bitrate = bus_match.groups()
            bus_config.append({
                'bus_no': int(bus_no),
                'name': name,
                'connection': conn,
                'protocol': proto,
                'bitrate': bitrate
            })
            continue
        
        # 匹配CAN报文数据
        msg_match = msg_pattern.match(line)
        if msg_match:
            seq, time_offset, bus_no, msg_type, can_id, dlc, data = msg_match.groups()
            can_messages.append({
                'seq': int(seq),
                'time_offset_ms': float(time_offset),
                'bus_no': int(bus_no),
                'trans_type': msg_type,
                'can_id_hex': can_id,
                'data_length': int(dlc),
                'data_bytes_hex': data.strip()
            })

# 直接转换为pandas DataFrame,可直接用于后续分析
bus_config_df = pd.DataFrame(bus_config)
can_msg_df = pd.DataFrame(can_messages)

3 可选优化方向

如果需要进一步提升易用性,可以做以下扩展:

  • 将data_bytes_hex字段按空格拆分,生成byte_0到byte_7的独立列,方便后续按位解析信号值
  • 针对J1939协议的ID,拆分出PGN、优先级、源地址、目标地址等专有字段
  • 把头部的开始时间戳转换为标准datetime格式,和时间偏移字段合并得到每条报文的绝对采集时间

内容的提问来源于stack exchange,提问作者Iceberg_Slim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 16:15:03