You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将XML格式数据集导入Brightway2/ActivityBrowser数据库?

导入PROBAS XML到Brightway2数据库的Python实现方案

一、Brightway2数据库的必填格式

Brightway2接受的数据库是字典组成的列表,每个字典对应一个活动(过程),核心必填字段如下:

  • name: 活动名称
  • unit: 活动的功能单位
  • exchanges: 交换项列表,每个交换项是字典,需包含:
    • name: 交换的产品/服务名称
    • amount: 交换的数值(需转为浮点数)
    • unit: 交换的单位
    • type: 交换类型("production"对应产出、"technosphere"对应技术输入、"biosphere"对应环境排放)
  • 可选常用字段:reference product(基准产品)、location(地理位置)、categories(分类,元组格式)

二、修复PROBAS XML的解析逻辑

你之前的递归解析会因为XML中重复标签(比如多个<exchange>的子标签)导致数据覆盖,无法正确提取交换项。针对PROBAS的XML结构,调整解析逻辑如下:

import xml.etree.ElementTree as ET
import pandas as pd

def parse_probas_xml(xml_path):
    tree = ET.parse(xml_path)
    root = tree.getroot()
    processes = []

    # 遍历每个过程节点
    for process in root.findall('.//process'):
        process_data = {
            'name': process.find('.//name').text.strip(),
            'unit': process.find('.//unit').text.strip(),
            'location': process.find('.//location').text.strip() if process.find('.//location') else 'GLO',
            'exchanges': []
        }

        # 提取所有交换项
        for exchange in process.findall('.//exchange'):
            exchange_type = exchange.find('.//type').text.strip()
            exchange_data = {
                'name': exchange.find('.//name').text.strip(),
                'amount': float(exchange.find('.//amount').text),
                'unit': exchange.find('.//unit').text.strip(),
                # 映射PROBAS类型到Brightway2标准类型
                'type': 'production' if exchange_type == 'output' else 'technosphere'
            }
            # 处理环境排放
            if exchange_type == 'emission':
                exchange_data['type'] = 'biosphere'
                exchange_data['categories'] = (exchange.find('.//category').text.strip(),) if exchange.find('.//category') else ('air',)
            
            process_data['exchanges'].append(exchange_data)
        
        processes.append(process_data)
    
    return processes

# 解析XML文件
brightway_data = parse_probas_xml('process.xml')
# 可选:转成DataFrame查看结构
df = pd.json_normalize(brightway_data, record_path='exchanges', meta=['name', 'unit', 'location'])

三、将数据写入Brightway2数据库

完成XML解析后,按以下步骤写入Brightway2:

import brightway2 as bw

# 1. 设置当前项目
bw.projects.set_current('PROBAS_Import_Project')  # 替换为你的项目名称

# 2. 初始化基础数据库(首次运行需执行)
# bw.bw2setup()

# 3. 创建并写入自定义数据库
db_name = 'PROBAS_Database'
if db_name not in bw.databases:
    db = bw.Database(db_name)
    db.write(brightway_data)
    print(f"数据库 {db_name} 导入成功,共 {len(db)} 个活动")
else:
    print(f"数据库 {db_name} 已存在,请先删除或重命名")

# 4. 验证导入结果
first_activity = next(bw.Database(db_name))
print("第一个活动详情:", first_activity.as_dict())

关键注意事项

  • PROBAS XML的标签命名可能存在差异,建议先打印XML节点结构(如ET.dump(root)),再调整解析逻辑中的节点路径。
  • 若XML中存在嵌套分类信息,需额外解析并转为Brightway2要求的元组格式(如('air', 'particulate matter'))。
  • 导入前确保Brightway2已正确安装,首次运行需执行bw.bw2setup()完成基础数据库配置。

内容的提问来源于stack exchange,提问作者squareK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 22:12:39