启用递归选项时Solr post.jar报错:content is not allowed in prolog
Environment Details
- OS: Windows Server 2012 R2
- Java Version:
1.8.0_171 - Solr Version: 7.3.0
Problem Description
I'm evaluating Solr and trying to crawl a local website using the post.jar tool with the recursive option enabled. When I run the following command:
java -Dauto=yes -Dc=testcore -Ddata=web -Drecursive=2 -Ddelay=10 -jar post.jar http://localhost/
I get the following error stack trace:
SimplePostTool version 5.0.0 Posting web pages to Solr url http://localhost:8983/solr/testcore/update/extract Entering auto mode. Indexing pages with content-types corresponding to file endings xml,json,jsonl,csv,pdf,doc,docx,ppt,pptx,xls,xlsx,odt,odp,ods,ott,otp,ots,rtf,htm,html,txt,log Entering recursive mode, depth=2, delay=10s Entering crawl at level 0 (1 links total, 1 new) POSTed web resource http://localhost/ (depth: 0) [Fatal Error] :1:1: Content is not allowed in prolog. Exception in thread "main" java.lang.RuntimeException: org.xml.sax.SAXParseException; lineNumber: 1; columnNumber: 1; Content is not allowed in prolog. at org.apache.solr.util.SimplePostTool$PageFetcher.getLinksFromWebPage(SimplePostTool.java:1252) at org.apache.solr.util.SimplePostTool.webCrawl(SimplePostTool.java:616) at org.apache.solr.util.SimplePostTool.postWebPages(SimplePostTool.java:563) at org.apache.solr.util.SimplePostTool.doWebMode(SimplePostTool.java:365) at org.apache.solr.util.SimplePostTool.execute(SimplePostTool.java:187) at org.apache.solr.util.SimplePostTool.main(SimplePostTool.java:172) Caused by: org.xml.sax.SAXParseException; lineNumber: 1; columnNumber: 1; Content is not allowed in prolog. at com.sun.org.apache.xerces.internal.parsers.DOMParser.parse(Unknown Source) at com.sun.org.apache.xerces.internal.jaxp.DocumentBuilderImpl.parse(Unknown Source) at javax.xml.parsers.DocumentBuilder.parse(Unknown Source) at org.apache.solr.util.SimplePostTool.makeDom(SimplePostTool.java:1061) at org.apache.solr.util.SimplePostTool$PageFetcher.getLinksFromWebPage(SimplePostTool.java:1232) ... 5 more
When I disable the recursive option (-Drecursive=2 removed), I can manually index all individual links (files and pages) from http://localhost/ without issues, so I don't think there are files or links with special characters causing this. I've searched for solutions without luck and would appreciate any help.
Answer
This error usually pops up because the SimplePostTool is trying to parse your HTML page as XML, which doesn't work since HTML isn't strict XML. Here are a few practical fixes to try:
Force HTML content type parsing
Add the-DcontentType=text/htmlparameter to your command. This tells the tool to use an HTML parser instead of an XML parser when extracting links during recursion, which should fix the parsing error:java -Dauto=yes -Dc=testcore -Ddata=web -Drecursive=2 -Ddelay=10 -DcontentType=text/html -jar post.jar http://localhost/Check for invalid leading characters in your page
Sometimes invisible characters (like a UTF-8 BOM) or extra whitespace before the opening<html>tag can trigger this XML parsing error. Right-click your local page, view its source, and make sure there's nothing before the first valid HTML tag.Validate your server's Content-Type header
Ensure your local web server sends the correctContent-Type: text/htmlheader for HTML pages. If it returns a generic or incorrect header (likeapplication/xml), the tool might misinterpret the content and try to parse it as XML.Upgrade to a newer Solr version
Solr 7.3.0 is pretty old (released in 2018), and there could be bugs in the recursive crawling logic that have been fixed in later releases. Since you're on Java 8, you can safely upgrade to a newer 7.x version (like 7.7.3) or even an 8.x release (most 8.x versions still support Java 8).
The first fix is the most likely to resolve your issue immediately—it directly addresses the core problem of the tool using the wrong parser during recursion.
内容的提问来源于stack exchange,提问作者EliudM

