请求提供Wikidata转Gremlin格式示例代码及AWS Neptune加载方案
将Wikidata转储转换为Gremlin CSV(适配AWS Neptune)
核心转换逻辑
Wikidata转储以JSON格式为主,核心结构包含实体(Q前缀ID)、属性(P前缀ID)、实体间关系。要适配Neptune的Gremlin CSV格式,需拆分为两类文件:
- 顶点CSV:存储实体,字段至少包含
~id(顶点唯一标识)、~label(顶点类型),附加自定义属性(如实体标签、描述) - 边CSV:存储实体间关系,字段至少包含
~id(边唯一标识)、~label(关系类型)、~from(源顶点ID)、~to(目标顶点ID)
Python示例代码
以下代码针对Wikidata的JSON片段进行转换,生成符合Neptune要求的CSV文件:
import json import csv from uuid import uuid4 def process_wikidata_dump(input_file, vertex_output, edge_output): # 初始化CSV写入器 vertex_writer = csv.writer(open(vertex_output, 'w', newline='', encoding='utf-8')) edge_writer = csv.writer(open(edge_output, 'w', newline='', encoding='utf-8')) # 写入CSV表头 vertex_writer.writerow(['~id', '~label', 'name:String', 'description:String']) edge_writer.writerow(['~id', '~label', '~from', '~to']) with open(input_file, 'r', encoding='utf-8') as f: for line in f: # 跳过Wikidata转储的首尾JSON数组标记行 if line.startswith('[') or line.startswith(']'): continue # 处理行尾多余逗号 if line.endswith(',\n'): line = line[:-2] data = json.loads(line) entity_id = data['id'] # 提取实体标签(优先取英文) label = data.get('labels', {}).get('en', {}).get('value', entity_id) # 提取实体描述(优先取英文) description = data.get('descriptions', {}).get('en', {}).get('value', '') # 写入顶点数据 vertex_writer.writerow([entity_id, 'Entity', label, description]) # 处理实体间指向其他实体的关系 for prop_id, claims in data.get('claims', {}).items(): for claim in claims: target = claim.get('mainsnak', {}).get('datavalue', {}).get('value', {}) if isinstance(target, dict) and 'id' in target: target_id = target['id'] edge_id = str(uuid4()) # 边标签用属性ID,也可替换为属性的英文名称 edge_writer.writerow([edge_id, prop_id, entity_id, target_id]) # 示例调用 if __name__ == '__main__': process_wikidata_dump( input_file='wikidata_sample.json', vertex_output='neptune_vertices.csv', edge_output='neptune_edges.csv' )
AWS Neptune加载注意事项
- 将生成的CSV文件上传至AWS S3桶,确保Neptune集群拥有该桶的访问权限
- 使用Neptune Bulk Loader加载数据,示例Gremlin命令:
g.withSideEffect('Neptune#load', {'source':'s3://your-bucket/path/neptune_vertices.csv', 'format':'csv', 'region':'us-east-1'}) g.withSideEffect('Neptune#load', {'source':'s3://your-bucket/path/neptune_edges.csv', 'format':'csv', 'region':'us-east-1'}) - 明确标注CSV属性类型(如
:String、:Int),避免Neptune自动推断数据类型出错 - 处理大规模转储时建议分块操作,可结合Spark等分布式框架提升处理效率
内容的提问来源于stack exchange,提问作者MAK
相关产品推荐
相关产品推荐

