You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu环境下如何用Python横向合并GridDB容器中的无共同列表格

在Ubuntu上用Python实现GridDB跨容器表横向拼接的方案

完全可以用Python实现这个需求,核心思路是通过GridDB Python客户端读取两个容器的表数据,用pandas做横向拼接补全NULL,再写入新的GridDB容器。

步骤1:安装依赖

先确保Ubuntu上装了GridDB Python客户端和pandas:

pip install griddb-python pandas

步骤2:完整实现代码

import griddb_python as griddb
import pandas as pd

# 两个GridDB容器的连接配置
# 容器1(存储table_1)的连接信息
container1_config = {
    "host": "your_container1_host",
    "port": 31810,
    "cluster_name": "your_cluster1_name",
    "username": "admin",
    "password": "admin"
}

# 容器2(存储table_2)的连接信息
container2_config = {
    "host": "your_container2_host",
    "port": 31810,
    "cluster_name": "your_cluster2_name",
    "username": "admin",
    "password": "admin"
}

# 新容器的配置(存储合并后的表)
new_container_config = {
    "host": "your_new_container_host",
    "port": 31810,
    "cluster_name": "your_new_cluster_name",
    "username": "admin",
    "password": "admin"
}
new_container_name = "merged_table_container"

def read_griddb_table(config, container_name):
    """读取GridDB容器中的表数据,返回DataFrame"""
    factory = griddb.StoreFactory.get_instance()
    try:
        store = factory.get_store(
            host=config["host"],
            port=config["port"],
            cluster_name=config["cluster_name"],
            username=config["username"],
            password=config["password"]
        )
        container = store.get_container(container_name)
        # 获取所有行数据
        query = container.query("select *")
        rs = query.fetch(False)
        # 获取列名
        column_names = [col.get_name() for col in container.get_column_info()]
        # 转换为DataFrame
        data = []
        while rs.has_next():
            data.append(rs.next())
        df = pd.DataFrame(data, columns=column_names)
        return df
    except Exception as e:
        print(f"读取表失败: {e}")
        raise

def create_griddb_container(config, container_name, column_info):
    """创建新的GridDB容器"""
    factory = griddb.StoreFactory.get_instance()
    try:
        store = factory.get_store(
            host=config["host"],
            port=config["port"],
            cluster_name=config["cluster_name"],
            username=config["username"],
            password=config["password"]
        )
        # 检查容器是否存在,存在则删除(可选)
        if store.container_exists(container_name):
            store.drop_container(container_name)
        # 定义容器 schema
        container_info = griddb.ContainerInfo(
            container_name,
            column_info,
            griddb.ContainerType.COLLECTION,
            True
        )
        container = store.put_container(container_info)
        return container
    except Exception as e:
        print(f"创建容器失败: {e}")
        raise

def write_df_to_griddb(df, container):
    """将DataFrame写入GridDB容器"""
    try:
        # 将DataFrame转换为tuple列表
        data_tuples = [tuple(row) for row in df.to_numpy()]
        container.multi_put(data_tuples)
        print("数据写入成功")
    except Exception as e:
        print(f"写入数据失败: {e}")
        raise

if __name__ == "__main__":
    # 1. 读取两个表的数据
    df1 = read_griddb_table(container1_config, "table_1")
    df2 = read_griddb_table(container2_config, "table_2")
    
    # 2. 横向拼接,按行索引对齐,不足的行填充NULL
    merged_df = pd.concat([df1, df2], axis=1, sort=False)
    
    # 3. 定义新容器的列信息(需要匹配merged_df的列和数据类型)
    # 示例:根据df1和df2的列信息组合,这里需要根据实际数据类型调整
    column_info = []
    # 添加table_1的列
    for col in df1.columns:
        # 这里需要根据实际数据类型对应GridDB的类型,比如STRING, INTEGER等
        # 示例默认用STRING,实际要根据你的表结构修改
        column_info.append(griddb.ColumnInfo(col, griddb.Type.STRING))
    # 添加table_2的列
    for col in df2.columns:
        column_info.append(griddb.ColumnInfo(col, griddb.Type.STRING))
    
    # 4. 创建新容器并写入数据
    new_container = create_griddb_container(new_container_config, new_container_name, column_info)
    write_df_to_griddb(merged_df, new_container)

关键说明

  • 数据读取:通过GridDB的query方法获取全表数据,转换为pandas DataFrame方便处理
  • 横向拼接:用pd.concat([df1, df2], axis=1)实现按行索引横向合并,pandas会自动给行数少的表补NaN(对应GridDB的NULL)
  • 容器创建:需要准确定义合并后表的列信息,数据类型要和原表一致(比如原表是INTEGER就用griddb.Type.INTEGER,STRING用griddb.Type.STRING)
  • 数据写入:将DataFrame转换为tuple列表,用multi_put批量写入GridDB,效率更高

注意事项

  • 确保两个GridDB容器的网络可以被Ubuntu机器访问,端口(默认31810)开放
  • 如果表数据量很大,建议分批读取和写入,避免内存溢出
  • 数据类型要严格对应,否则写入时会报错

内容的提问来源于stack exchange,提问作者Jessica Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 02:15:30