You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取类HTML标签内的文本数据并转换为Pandas DataFrame

最优实现方案

核心思路是优先用XML解析器处理结构化的标签内容,避免手写正则匹配每个字段带来的维护问题,同时自动处理嵌套节点和JSON格式的Address字段,步骤如下:

  1. 用正则从日志行中提取两种目标标签包裹的XML片段
  2. 用Python内置XML库解析XML节点,拉平所有层级的字段
  3. 解析Address字段的JSON字符串,将地址明细也展开为独立列
  4. 所有解析结果收集为字典列表后直接转为Pandas DataFrame
import re
import xml.etree.ElementTree as ET
import json
import pandas as pd
from collections import OrderedDict

# 示例数据,可替换为实际数据源
raw_data = [
    OrderedDict([('_raw', '2021-11-08 08:58:23,832 [42] INFO  FiservLog.stdlog - <NetworkPayeeAddManager><TenantId>13744</TenantId><UserId>999176993878</UserId><SourceMethodName>LogInfoSecure</SourceMethodName><SourceLineNumber>234</SourceLineNumber><Message>NetworkPayee was added successfully</Message><Timestamp>2021-11-08T13:58:23.831628Z</Timestamp><Exception /><AdditionalInformation><SessionId>F7E65ED4D8C74E6699C62F23ECF5D000200TWNQ9X1AA1754513234A6367FEE06</SessionId><Timestamp>11/8/2021 1:58:23 PM</Timestamp><CorrelationId>2461b5d9839a46739e9a3e918ca0681b-01</CorrelationId><PayeeName>Louisville fire brick</PayeeName><Address>{"Address1":"Po 9229","Address2":null,"City":"Louisville","State":"KY","Zip5":"40209","Zip4":null,"Zip2":null}</Address><PayeeType>UnManagedPayee</PayeeType><AccountNumber>XX2222</AccountNumber></AdditionalInformation></NetworkPayeeAddManager>')]),
    OrderedDict([('_raw', '2021-11-08 08:58:24,783 [105] INFO  FiservLog.stdlog - <PayeeAddManager><TenantId>DI737</TenantId><UserId>344801483</UserId> <SourceMethodName>LogInfoSecure</SourceMethodName><SourceLineNumber>234</SourceLineNumber><Message>Payee was added successfully</Message><Timestamp>2021-11-08T13:58:24.7831103Z</Timestamp><Exception /><AdditionalInformation><SessionId>7FC6442718864CE4838E50B026C8D0A0000TWNXSV1721BE0D804F295706DD39E</SessionId><Timestamp>11/8/2021 1:58:24 PM</Timestamp><CorrelationId>ab33b59c-756e-4144-ad62-6f0afadbe8eb</CorrelationId><PayeeName>Gail Nezworski</PayeeName><Address>{"Address1":"2280 S 460 E","Address2":null,"City":"LaGrange","State":"IN","Zip5":"46761","Zip4":null,"Zip2":null}</Address><PayeeType>UnManagedPayee</PayeeType><AccountNumber>XXXXX1888</AccountNumber></AdditionalInformation></PayeeAddManager>')])
]

# 匹配两种目标标签的正则
pattern = re.compile(r'<(NetworkPayeeAddManager|PayeeAddManager)>(.*?)</\1>', re.DOTALL)
result_list = []

for item in raw_data:
    raw_log = item['_raw']
    match = pattern.search(raw_log)
    if not match:
        continue
    # 构造合法XML结构便于解析
    xml_content = f"<root>{match.group(2)}</root>"
    root = ET.fromstring(xml_content)
    row = {}
    # 解析根节点下的直接字段
    for child in root:
        if child.tag == 'AdditionalInformation':
            # 解析附加信息的嵌套字段
            for sub_child in child:
                if sub_child.tag == 'Timestamp':
                    # 重名区分本地时间和UTC时间
                    row['Timestamp_Local'] = sub_child.text.strip() if sub_child.text else None
                elif sub_child.tag == 'Address':
                    # 解析地址JSON展开为独立列
                    if sub_child.text:
                        addr_dict = json.loads(sub_child.text)
                        row.update(addr_dict)
                        # 不需要保留原始Address字段可注释下一行
                        row['Address'] = sub_child.text
                else:
                    row[sub_child.tag] = sub_child.text.strip() if sub_child.text else None
        elif child.tag == 'Timestamp':
            row['Timestamp_UTC'] = child.text.strip() if child.text else None
        else:
            row[child.tag] = child.text.strip() if child.text else None
    result_list.append(row)

# 转为DataFrame
df = pd.DataFrame(result_list)

实现优势

  • 用XML解析比正则匹配每个字段更稳定,后续字段增减不需要修改匹配规则
  • 自动处理了重名的Timestamp字段,分别标记为UTC时间和本地时间
  • Address字段自动展开为Address1、City、State等独立列,也可按需保留原始字符串
  • 空节点比如Exception会自动赋值为None,适配Pandas的空值处理规范

内容的提问来源于stack exchange,提问作者bergen288

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 11:24:07