You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

启用递归选项时Solr post.jar报错:content is not allowed in prolog

Solr post.jar Recursive Crawl Fails with "Content is not allowed in prolog" Error

Environment Details

  • OS: Windows Server 2012 R2
  • Java Version: 1.8.0_171
  • Solr Version: 7.3.0

Problem Description

I'm evaluating Solr and trying to crawl a local website using the post.jar tool with the recursive option enabled. When I run the following command:

java -Dauto=yes -Dc=testcore -Ddata=web -Drecursive=2 -Ddelay=10 -jar post.jar http://localhost/

I get the following error stack trace:

SimplePostTool version 5.0.0
Posting web pages to Solr url http://localhost:8983/solr/testcore/update/extract
Entering auto mode. Indexing pages with content-types corresponding to file endings xml,json,jsonl,csv,pdf,doc,docx,ppt,pptx,xls,xlsx,odt,odp,ods,ott,otp,ots,rtf,htm,html,txt,log
Entering recursive mode, depth=2, delay=10s
Entering crawl at level 0 (1 links total, 1 new)
POSTed web resource http://localhost/ (depth: 0)
[Fatal Error] :1:1: Content is not allowed in prolog.
Exception in thread "main" java.lang.RuntimeException: org.xml.sax.SAXParseException; lineNumber: 1; columnNumber: 1; Content is not allowed in prolog.
	at org.apache.solr.util.SimplePostTool$PageFetcher.getLinksFromWebPage(SimplePostTool.java:1252)
	at org.apache.solr.util.SimplePostTool.webCrawl(SimplePostTool.java:616)
	at org.apache.solr.util.SimplePostTool.postWebPages(SimplePostTool.java:563)
	at org.apache.solr.util.SimplePostTool.doWebMode(SimplePostTool.java:365)
	at org.apache.solr.util.SimplePostTool.execute(SimplePostTool.java:187)
	at org.apache.solr.util.SimplePostTool.main(SimplePostTool.java:172)
Caused by: org.xml.sax.SAXParseException; lineNumber: 1; columnNumber: 1; Content is not allowed in prolog.
	at com.sun.org.apache.xerces.internal.parsers.DOMParser.parse(Unknown Source)
	at com.sun.org.apache.xerces.internal.jaxp.DocumentBuilderImpl.parse(Unknown Source)
	at javax.xml.parsers.DocumentBuilder.parse(Unknown Source)
	at org.apache.solr.util.SimplePostTool.makeDom(SimplePostTool.java:1061)
	at org.apache.solr.util.SimplePostTool$PageFetcher.getLinksFromWebPage(SimplePostTool.java:1232)
	... 5 more

When I disable the recursive option (-Drecursive=2 removed), I can manually index all individual links (files and pages) from http://localhost/ without issues, so I don't think there are files or links with special characters causing this. I've searched for solutions without luck and would appreciate any help.


Answer

This error usually pops up because the SimplePostTool is trying to parse your HTML page as XML, which doesn't work since HTML isn't strict XML. Here are a few practical fixes to try:

  1. Force HTML content type parsing
    Add the -DcontentType=text/html parameter to your command. This tells the tool to use an HTML parser instead of an XML parser when extracting links during recursion, which should fix the parsing error:

    java -Dauto=yes -Dc=testcore -Ddata=web -Drecursive=2 -Ddelay=10 -DcontentType=text/html -jar post.jar http://localhost/
    
  2. Check for invalid leading characters in your page
    Sometimes invisible characters (like a UTF-8 BOM) or extra whitespace before the opening <html> tag can trigger this XML parsing error. Right-click your local page, view its source, and make sure there's nothing before the first valid HTML tag.

  3. Validate your server's Content-Type header
    Ensure your local web server sends the correct Content-Type: text/html header for HTML pages. If it returns a generic or incorrect header (like application/xml), the tool might misinterpret the content and try to parse it as XML.

  4. Upgrade to a newer Solr version
    Solr 7.3.0 is pretty old (released in 2018), and there could be bugs in the recursive crawling logic that have been fixed in later releases. Since you're on Java 8, you can safely upgrade to a newer 7.x version (like 7.7.3) or even an 8.x release (most 8.x versions still support Java 8).

The first fix is the most likely to resolve your issue immediately—it directly addresses the core problem of the tool using the wrong parser during recursion.

内容的提问来源于stack exchange,提问作者EliudM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:06:37