You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Neo4j数据库导出数据导入ArangoDB?是否有对应工具?

把Neo4j数据导入ArangoDB的解决方案

我来分享下处理这类需求的实际经验,分成手动转换导入和工具推荐两部分来说:

一、手动转换格式后用arangoimp导入

如果不想依赖第三方工具,手动调整CSV格式是最直接的方案——毕竟arangoimp本身性能不错,只是和Neo4j导出的格式不兼容,步骤如下:

  1. 从Neo4j导出CSV数据
    用Neo4j自带的neo4j-admin export命令导出节点和关系,示例命令:

    neo4j-admin export --database=neo4j --nodes=nodes.csv --relationships=relationships.csv --delimiter=,
    

    导出的nodes.csv会包含节点ID、标签和属性,relationships.csv则包含起始节点ID、结束节点ID、关系类型和属性。

  2. 转换顶点(节点)CSV格式
    ArangoDB的顶点CSV必须包含_key字段(作为顶点的唯一标识),你可以把Neo4j的:ID字段直接重命名为_key(注意转成字符串,因为ArangoDB的_key仅支持字符串类型),也可以添加_collection字段指定目标集合名。
    举个转换前后的例子:
    原Neo4j节点CSV:

    :ID,label:LABEL,name,age
    1,Person,Alice,30
    2,Person,Bob,25
    

    转换后:

    _key,label,name,age
    "1",Person,Alice,30
    "2",Person,Bob,25
    
  3. 转换边(关系)CSV格式
    ArangoDB的边CSV需要_from和_to字段,格式为[顶点集合名]/[顶点_key],关系类型可以保留为属性,也可以用来命名边集合。
    原Neo4j关系CSV:

    :START_ID,:END_ID,:TYPE,friend_since
    1,2,FRIEND,2020
    

    转换后(假设顶点集合叫nodes):

    _from,_to,type,friend_since
    nodes/1,nodes/2,FRIEND,2020
    
  4. 用arangoimp导入
    分别执行顶点和边的导入命令:

    # 导入顶点
    arangoimp --file nodes_converted.csv --type csv --collection nodes --create-collection true
    # 导入边(注意指定--create-collection-type edge)
    arangoimp --file edges_converted.csv --type csv --collection relationships --create-collection true --create-collection-type edge
    

二、可用的迁移工具推荐

如果数据量较大或者不想手动改格式,这些工具能帮你节省时间:

  • 自定义Python脚本
    这是最灵活的方式,适合有编程基础的开发者。用py2neo连接Neo4j读取数据,再用python-arango写入ArangoDB,示例脚本如下:

    from py2neo import Graph
    from arango import ArangoClient
    
    # 连接Neo4j
    neo4j_graph = Graph("bolt://localhost:7687", auth=("neo4j", "your_password"))
    # 连接ArangoDB
    arango_client = ArangoClient(hosts="http://localhost:8529")
    db = arango_client.db("your_db", username="root", password="your_password")
    
    # 创建顶点和边集合
    nodes_col = db.create_collection("nodes")
    edges_col = db.create_collection("relationships", edge=True)
    
    # 批量导入节点(避免内存溢出,每1000条提交一次)
    nodes_batch = []
    for node in neo4j_graph.run("MATCH (n) RETURN n.id AS id, labels(n) AS labels, properties(n) AS props"):
        nodes_batch.append({
            "_key": str(node["id"]),
            "labels": node["labels"],
            **node["props"]
        })
        if len(nodes_batch) >= 1000:
            nodes_col.insert_many(nodes_batch)
            nodes_batch = []
    if nodes_batch:
        nodes_col.insert_many(nodes_batch)
    
    # 批量导入边
    edges_batch = []
    for rel in neo4j_graph.run("MATCH ()-[r]->() RETURN r.startNodeId AS start_id, r.endNodeId AS end_id, type(r) AS type, properties(r) AS props"):
        edges_batch.append({
            "_from": f"nodes/{rel['start_id']}",
            "_to": f"nodes/{rel['end_id']}",
            "type": rel["type"],
            **rel["props"]
        })
        if len(edges_batch) >= 1000:
            edges_col.insert_many(edges_batch)
            edges_batch = []
    if edges_batch:
        edges_col.insert_many(edges_batch)
    
  • 社区开源工具
    可以去GitHub上搜索neo4j-arangodb-importer这类项目,不过要注意检查工具是否兼容你当前使用的Neo4j和ArangoDB版本,避免出现兼容性问题。

注意事项

  • 确保Neo4j的节点ID转成ArangoDB的_key后是唯一的,避免导入冲突;
  • 大规模数据迁移时,务必分批次处理,防止内存占用过高;
  • 迁移完成后要验证数据完整性,比如对比节点数、边数,抽样检查属性是否正确。

内容的提问来源于stack exchange,提问作者user2727704

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:28:15