使用MLCP导出MarkLogic带查询过滤集合时遇XDMP-DOCROOTTEXT错误
MLCP导出MarkLogic数据时遭遇XDMP-DOCROOTTEXT错误
错误日志
WARNING: An illegal reflective access operation has occurred WARNING: Illegal reflective access by org.apache.hadoop.security.authentication.util.KerberosUtil (file:/opt/MarkLogic/mlcp-10.0.8.2/lib/hadoop-auth-2.7.2.jar) to method sun.security.krb5.Config.getInstance() WARNING: Please consider reporting this to the maintainers of org.apache.hadoop.security.authentication.util.KerberosUtil WARNING: Use --illegal-access=warn to enable warnings of further illegal reflective access operations WARNING: All illegal access operations will be denied in a future release 24/10/14 04:26:31 INFO contentpump.ContentPump: Job name: local_928393857_1 24/10/14 04:26:31 ERROR mapreduce.MarkLogicInputFormat: com.marklogic.xcc.exceptions.XQueryException: XDMP-DOCROOTTEXT: xdmp:unquote("collection_name_SerializedQuery.txt") -- Invalid root text "collection_name_SerializedQuery.txt" at line 1 [Session: user=username, cb={default} [ContentSource: user=username, cb={none} [provider: SSLconn address=hostname/x.x.x.x:****, pool=1/64]]] [Client: XCC/10.0-8, Server: XDBC/10.0-11] on line 1 expr: xdmp:unquote("collection_name_SerializedQuery.txt"), in xdmp:eval("for $f in xdmp:forest-open-replica(xdmp:database-forests(xdmp:da...") in /MarkLogic/hadoop.xqy, on line 32 expr: xdmp:unquote("collection_name_SerializedQuery.txt"), in hadoop:get-splits("", "fn:collection("gps-temporal")", "cts:query(xdmp:unquote('collection_name_SerializedQuery.txt')/*)") on line 5 expr: xdmp:unquote("collection_name_SerializedQuery.txt") 24/10/14 04:26:31 ERROR mapreduce.MarkLogicInputFormat: Query: xquery version "1.0-ml"; fn:exists(xdmp:get-request-header('x-forwarded-for')); import module namespace hadoop = "http://marklogic.com/xdmp/hadoop" at "/MarkLogic/hadoop.xqy"; xdmp:host-name(xdmp:host()), hadoop:get-splits('', 'fn:collection("gps-temporal")',"cts:query(xdmp:unquote('collection_name_SerializedQuery.txt')/*)"), "REDACT",0,let $repf := fn:function-lookup(xs:QName('hadoop:get-splits-with-replica'),0) return if (exists($repf)) then $repf() else () ,0,"AUDIT", let $f := fn:function-lookup(xs:QName('xdmp:group-get-audit-event-type-enabled'), 2) return if (not(exists($f))) then () else let $group-id := xdmp:group() let $enabled-event := $f($group-id,("mlcp-copy-export-start", "mlcp-copy-export-finish")) let $mlcp-start-enabled := if ($enabled-event[1]) then "mlcp-copy-export-start" else () let $mlcp-finish-enabled := if ($enabled-event[2]) then "mlcp-copy-export-finish" else () return ($mlcp-start-enabled, $mlcp-finish-enabled) 24/10/14 04:26:31 ERROR contentpump.LocalJobRunner: Error getting input splits: 24/10/14 04:26:31 ERROR contentpump.LocalJobRunner: com.marklogic.xcc.exceptions.XQueryException: XDMP-DOCROOTTEXT: xdmp:unquote("collection_name_SerializedQuery.txt") -- Invalid root text "collection_name_SerializedQuery.txt" at line 1 [Session: user=sc799-sa, cb={default} [ContentSource: user=sc799-sa, cb={none} [provider: SSLconn address=mlg-gpi-uat.ntrs.com/10.33.132.50:8062, pool=1/64]]] [Client: XCC/10.0-8, Server: XDBC/10.0-11] on line 1 expr: xdmp:unquote("collection_name_SerializedQuery.txt"), in xdmp:eval("for $f in xdmp:forest-open-replica(xdmp:database-forests(xdmp:da...") in /MarkLogic/hadoop.xqy, on line 32 expr: xdmp:unquote("collection_name_SerializedQuery.txt"), in hadoop:get-splits("", "fn:collection("gps-temporal")", "cts:query(xdmp:unquote('collection_name_SerializedQuery.txt')/*)") on line 5 expr: xdmp:unquote("collection_name_SerializedQuery.txt")
导出任务配置(optionsFile.txt)
-host <hostname> -port **** -ssl true -username username -password ****** -collection_filter collection_name -query_filter collection_SerializedQuery.txt -output_file_path <path> -output_type archive -thread_count 8
查询过滤文件(collection_SerializedQuery.txt)
<?xml version="1.0" encoding="UTF-8"?> <query xmlns:cts="http://marklogic.com/cts"><cts:uri>/collection_name</cts:uri> <and-query> <collection-query> <collection>collection_name</collection> </collection-query> <json-property-scope-query> <property-name>metadata</property-name> <query> <json-property-value-query> <property-name>archivalDate</property-name> <value>2023-11-30</value> <value>2023-11-29</value> </json-property-value-query> </query> </json-property-scope-query> </and-query> </query>
问题原因
从错误日志可见,MLCP将-query_filter后的collection_SerializedQuery.txt直接作为字符串传入xdmp:unquote()函数,而非读取该文件的XML查询内容。xdmp:unquote()尝试将纯文本文件名解析为XML,因文件名不是合法XML结构,触发XDMP-DOCROOTTEXT错误。
核心问题是参数误用:-query_filter用于直接传入序列化的cts查询字符串,指定包含查询内容的文件路径需使用-query_filter_file参数。
解决方案
修改导出选项文件,将-query_filter替换为-query_filter_file,修改后的配置如下:
-host <hostname> -port **** -ssl true -username username -password ****** -collection_filter collection_name -query_filter_file collection_SerializedQuery.txt -output_file_path <path> -output_type archive -thread_count 8
重新执行导出命令:
sh mlcp.sh export -options_file optionsFile.txt
内容的提问来源于stack exchange,提问作者Sravan
相关产品推荐
相关产品推荐

