Gremlin Python加载CSV时出现org.apache.commons.csv.CSVFormat类解析错误
问题描述
使用Gremlin Python加载CSV文件时出现以下错误:
{'error': GremlinServerError('597: startup failed:\r\nScript5.groovy: 2: unable to resolve class org.apache.commons.csv.CSVFormat\n @ line 2, column 1.\n import org.apache.commons.csv.CSVFormat\n ^\n\n1 error\n')}
已在Gremlin控制台成功加载该CSV文件(包含顶点和边),Python端代码如下:
from gremlin_python.driver.driver_remote_connection import DriverRemoteConnection from gremlin_python.structure.graph import Graph from gremlin_python import statics from gremlin_python.process.graph_traversal import __ from gremlin_python.process.strategies import * from gremlin_python.process.traversal import * import sys from gremlin_python.driver.aiohttp.transport import AiohttpTransport # Path to our graph (this assumes a locally running Gremlin Server) # Note how the path is a Web Socket (ws) connection. endpoint = 'ws://localhost:8182/gremlin' # Obtain a graph traversal source using a remote connection graph=Graph() connection = DriverRemoteConnection(endpoint,'g', transport_factory=lambda:AiohttpTransport(call_from_event_loop=True)) g = graph.traversal().withRemote(connection) %%graph_notebook_config { "host": "localhost", "port": 8182, "ssl": false, "gremlin": { "traversal_source": "g" } } %%gremlin g = TinkerGraph.open().traversal() import org.apache.commons.csv.CSVFormat fileReader = new FileReader('C:/airports.csv') records = CSVFormat.RFC4180.withFirstRecordAsHeader().parse(fileReader);[] records.each{ c=it.get('code'); d=it.get('desc'); println(it) g.V().has('code', c).fold().coalesce( unfold(), addV('airport').property('code',c).property('desc',d) ).iterate() } g.V().count()
在Gremlin控制台中通过以下命令配置服务器:
:remote connect tinkerpop.server conf/remote.yaml
CSV文件包含两列,示例数据:code:TXF,desc:"9 de Maio - Teixeira de Freitas Airport"
解决方案
1. 补充Gremlin Server依赖包
错误提示找不到org.apache.commons.csv.CSVFormat类,说明Gremlin Server的classpath中缺少Apache Commons CSV依赖:
- 下载与Gremlin Server版本兼容的Apache Commons CSV jar包
- 将jar包复制到Gremlin Server安装目录下的
lib文件夹 - 重启Gremlin Server
2. 修正Gremlin脚本逻辑
Python代码中%%gremlin块的脚本存在两处问题:
- 重新创建了TinkerGraph实例,未使用已配置的遍历源
g - 脚本末尾多余的
;[]会导致执行异常
修改后的脚本:
import org.apache.commons.csv.CSVFormat fileReader = new FileReader('C:/airports.csv') records = CSVFormat.RFC4180.withFirstRecordAsHeader().parse(fileReader) records.each{ c=it.get('code'); d=it.get('desc'); g.V().has('code', c).fold().coalesce( unfold(), addV('airport').property('code',c).property('desc',d) ).iterate() } g.V().count()
3. 替代方案:Python本地读取CSV后批量插入
若不想修改Gremlin Server依赖,可在Python端直接读取CSV,再通过Gremlin Python发送遍历请求插入数据:
import csv # 读取本地CSV文件 with open('C:/airports.csv', 'r', encoding='utf-8') as f: reader = csv.DictReader(f) for row in reader: code = row['code'] desc = row['desc'] # 执行插入逻辑 g.V().has('code', code).fold().coalesce( __.unfold(), __.addV('airport').property('code', code).property('desc', desc) ).iterate() # 统计顶点数量 print(g.V().count().next())
内容的提问来源于stack exchange,提问作者Aureon
相关产品推荐
相关产品推荐

