You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取长XML文件提取字段至DataFrame遇空值问题求助

解决XML数据提取到DataFrame的问题

问题根源

你的代码未定位到正确的目标节点,且没有完成字段值提取、数据收集的完整流程,最终导致DataFrame为空。

修正后的代码

import xml.etree.ElementTree as ET
import pandas as pd

# 解析XML文件
tree = ET.parse(r'path')
root = tree.getroot()

# 定义DataFrame列名
df_cols = ["idInstrument", "shortname", "quantita"]
rows = []

# 精准定位目标节点:遍历<RiskFullInfo>下<stockAmounts>中的所有<stockAmount>子节点
# 注:若XML包含命名空间,需按下方说明额外处理
for stock_amount in root.findall('.//stockAmounts/stockAmount'):
    # 提取字段文本值
    instrument_id = stock_amount.findtext('idInstrument')
    short_name = stock_amount.findtext('shortname')
    quantita_val = stock_amount.findtext('quantita')
    
    # 将单条数据存入行列表
    rows.append({
        df_cols[0]: instrument_id,
        df_cols[1]: short_name,
        df_cols[2]: quantita_val
    })

# 转换为DataFrame
df = pd.DataFrame(rows, columns=df_cols)
print(df)

关键说明

  • 节点定位:使用findall('.//stockAmounts/stockAmount')直接锁定所有目标子节点,避免无效遍历root的顶层子节点。
  • 字段提取:findtext()可直接获取元素的文本内容,比find()返回元素对象后再取.text更高效。
  • 数据收集:必须将提取到的字段值组装成字典并添加到rows列表,否则无法生成有效DataFrame。
  • 命名空间兼容:如果XML带有命名空间(如<ns:RiskFullInfo>),需先定义命名空间映射再提取:
    ns = {'ns': 'http://your-namespace-url'}
    for stock_amount in root.findall('.//ns:stockAmounts/ns:stockAmount', namespaces=ns):
        instrument_id = stock_amount.findtext('ns:idInstrument', namespaces=ns)
        # 其余字段提取需同步加上命名空间前缀
    

内容的提问来源于stack exchange,提问作者CristinaK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 00:07:22