You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则逐行提取文本数据并生成指定格式的Excel表格?

解决XML数据提取并生成Excel的问题

以下是修改后的代码,严格按照指定正则规则提取字段,并生成目标Excel表格:

import re
import pandas as pd

# 初始化存储提取结果的列表
extracted_data = []

# 读取test.txt文件并逐行处理
with open('test.txt', 'r', encoding='utf-8') as txt_file:
    for line in txt_file:
        cleaned_line = line.strip()
        # 跳过空行
        if not cleaned_line:
            continue
        
        # 提取Request Type
        req_type_result = re.search(r'<(.*)\s+xmlns', cleaned_line)
        req_type = req_type_result.group(1) if req_type_result else '未匹配到'
        
        # 提取Channel
        channel_result = re.search(r'<Channel>(.*)</Channel>', cleaned_line)
        channel = channel_result.group(1) if channel_result else '未匹配到'
        
        # 提取AG参数值,re.DOTALL允许.匹配换行符,处理跨多行的XML标签
        ag_result = re.search(r'<Param Name="AG">.*?<Value>(.*?)</Value>', cleaned_line, re.DOTALL)
        ag_value = ag_result.group(1) if ag_result else '未匹配到'
        
        # 将当前行的提取结果加入列表
        extracted_data.append({
            'Request Type': req_type,
            'Channel': channel,
            'AG': ag_value
        })

# 转换为DataFrame并生成Excel
result_df = pd.DataFrame(extracted_data)
result_df.to_excel('extracted_result.xlsx', index=False)

关键说明

  • 加入re.DOTALL参数处理AG字段可能跨多行的XML结构,避免因标签换行导致匹配失败
  • 增加空行跳过逻辑,过滤无效行
  • 每个正则匹配都做了空值判断,防止匹配失败抛出异常
  • 指定文件编码为utf-8,避免读取中文XML时出现乱码

内容的提问来源于stack exchange,提问作者Rejoy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 01:05:34