如何在数据库中插入非字符串对象作为键?附MongoDB报错解决及替代方案
首先,你遇到的InvalidDocument错误确实是MongoDB的硬性限制——它要求文档的键必须是字符串类型,而你用了元组('good','beneficial')作为edge_attr的键,这就触发了报错。先给你一个快速解决MongoDB问题的方法,再针对你的图数据存储需求推荐更合适的方案。
一、快速修复MongoDB的键错误
把元组键转换成字符串就能解决问题,这里有两种常用方式:
方式1:用分隔符拼接元组
把元组的两个元素用下划线(或其他不会冲突的字符)拼接成字符串键,后续读取时可以再拆分还原:
import pymongo from pymongo import MongoClient client = MongoClient() db = client['glot'] collection = db['usr_history'] # 把元组键转成字符串键 edge_attr = { "_".join(edge): desc for edge, desc in [ (('good','beneficial'), "good is beneficial"), (('beneficial','supportive'), "beneficial means something supportive") ] } usr_history = { "nodes": ['good', 'beneficial', 'supportive'], "edges": [('good', 'beneficial'), ('beneficial', 'supportive')], "rephrase": edge_attr } collection.insert_one(usr_history)
方式2:JSON序列化元组
如果担心拆分时出现冲突(比如顶点名称里有下划线),可以用JSON把元组序列化成字符串,后续读取时用json.loads还原:
import json import pymongo from pymongo import MongoClient client = MongoClient() db = client['glot'] collection = db['usr_history'] edge_attr = { json.dumps(edge): desc for edge, desc in [ (('good','beneficial'), "good is beneficial"), (('beneficial','supportive'), "beneficial means something supportive") ] } usr_history = { "nodes": ['good', 'beneficial', 'supportive'], "edges": [('good', 'beneficial'), ('beneficial', 'supportive')], "rephrase": edge_attr } collection.insert_one(usr_history) # 读取时还原元组键 doc = collection.find_one() recovered_edge_attr = { json.loads(key): value for key, value in doc['rephrase'].items() }
二、更适合图数据的存储方案
你的核心需求是存储顶点、边、边属性这类图结构数据,还要用Python做分析,MongoDB虽然能改,但并不是最优选择。这里给你几个更匹配的方案:
1. 原生图数据库(推荐)
这类数据库专门为图数据设计,能高效处理顶点关系的遍历、查询,Python生态也很完善:
Neo4j
最流行的开源图数据库,用Cypher语言操作图数据非常直观,官方Python驱动neo4j很易用:
from neo4j import GraphDatabase # 连接数据库(替换成你的账号密码) driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "your_password")) def save_graph_data(tx, nodes, edges, edge_attrs): # 创建所有顶点 for node_name in nodes: tx.run("CREATE (n:Term {name: $name})", name=node_name) # 创建边并添加属性 for start, end in edges: desc = edge_attrs.get((start, end), "") tx.run( "MATCH (a:Term {name: $start}), (b:Term {name: $end}) " "CREATE (a)-[r:RELATION {description: $desc}]->(b)", start=start, end=end, desc=desc ) # 执行数据插入 with driver.session() as session: session.execute_write( save_graph_data, nodes=['good', 'beneficial', 'supportive'], edges=[('good', 'beneficial'), ('beneficial', 'supportive')], edge_attrs={('good','beneficial'):"good is beneficial", ('beneficial','supportive'):"beneficial means something supportive"} ) driver.close()
后续做图分析(比如找最短路径、关联顶点)时,Cypher语句会比MongoDB的查询简洁高效很多。
ArangoDB
支持多模型(文档、图、键值)的数据库,既保留文档数据库的灵活性,又有原生图数据库的能力,Python驱动python-arango可以轻松操作。
2. 关系型数据库(结构化存储)
如果习惯用传统数据库,PostgreSQL(或MySQL)也能很好地建模图数据,通过三个表来拆分存储:
nodes:存储顶点ID和名称edges:存储边ID、起点ID、终点IDedge_attributes:存储边ID和对应的属性
用SQLAlchemy(Python ORM框架)的示例:
from sqlalchemy import create_engine, Column, Integer, String, ForeignKey from sqlalchemy.ext.declarative import declarative_base from sqlalchemy.orm import sessionmaker Base = declarative_base() # 定义表结构 class Node(Base): __tablename__ = 'nodes' id = Column(Integer, primary_key=True) name = Column(String, unique=True, nullable=False) class Edge(Base): __tablename__ = 'edges' id = Column(Integer, primary_key=True) start_node_id = Column(Integer, ForeignKey('nodes.id'), nullable=False) end_node_id = Column(Integer, ForeignKey('nodes.id'), nullable=False) class EdgeAttribute(Base): __tablename__ = 'edge_attributes' edge_id = Column(Integer, ForeignKey('edges.id'), primary_key=True) description = Column(String) # 初始化数据库连接(替换成你的数据库信息) engine = create_engine('postgresql://username:password@localhost:5432/glot') Base.metadata.create_all(engine) Session = sessionmaker(bind=engine) session = Session() # 插入顶点 node_list = [Node(name='good'), Node(name='beneficial'), Node(name='supportive')] session.add_all(node_list) session.commit() # 映射顶点名称到ID node_id_map = {n.name: n.id for n in node_list} # 插入边 edge_list = [ Edge(start_node_id=node_id_map['good'], end_node_id=node_id_map['beneficial']), Edge(start_node_id=node_id_map['beneficial'], end_node_id=node_id_map['supportive']) ] session.add_all(edge_list) session.commit() # 插入边属性 attr_list = [ EdgeAttribute(edge_id=edge_list[0].id, description="good is beneficial"), EdgeAttribute(edge_id=edge_list[1].id, description="beneficial means something supportive") ] session.add_all(attr_list) session.commit() session.close()
这种方式结构清晰,适合需要做复杂统计分析的场景,配合PostgreSQL的扩展还能提升查询性能。
总结
- 继续用MongoDB:用字符串转换元组键即可解决报错,适合简单的图数据存储。
- 高效图分析:优先选Neo4j或ArangoDB这类原生图数据库。
- 结构化存储+统计:用PostgreSQL等关系型数据库建模。
内容的提问来源于stack exchange,提问作者snapper

