从指定URL读取XML遇解析错误,添加UTF-8编码仍未解决
Hey there, let's break down what's going on here and get your code working properly!
The Core Problem
First off, the root issue is that you’re using an XML parser to try and parse an HTML page. The URL http://www.bnb.bg/ returns regular HTML, not well-formed XML. HTML often has loose syntax (like unquoted attributes, unclosed tags, or elements like <meta> that skip closing tags) that strict XML parsers can’t handle—that’s exactly why you’re seeing the Open quote is expected for attribute "http-equiv" error. Adding UTF-8 encoding won’t fix this, because the problem isn’t character encoding—it’s that the content isn’t valid XML at all.
Bonus Bug in Your Code
On top of that, your code has a NullPointerException waiting to trigger:
Transformer transformer = null; transformer.setOutputProperty(OutputKeys.OMIT_XML_DECLARATION,"no");
You initialize transformer to null then immediately call a method on it. This line is also unnecessary—you already have the xform Transformer instance, so you should use that instead if you were to continue with XML parsing.
The Solution: Use an HTML Parser
Instead of forcing HTML through an XML parser, use a library built specifically for parsing messy web HTML. Jsoup is the standard choice for Java. Here’s how to rewrite your code to work correctly:
First, add Jsoup to your project. If you’re using Maven, add this dependency to your pom.xml:
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> </dependency>
Then, here’s the updated code:
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import java.io.IOException; public class CrawleyCraw { public static void main(String[] args) throws IOException { String urlString = "http://www.bnb.bg/"; // Fetch and parse the HTML page with Jsoup (handles messy syntax automatically) Document doc = Jsoup.connect(urlString).get(); // Print the pretty-printed HTML System.out.println(doc.html()); // Or convert to well-formed XHTML if you need XML-compatible output System.out.println(doc.outerHtml()); } }
Jsoup takes care of all the unstructured HTML quirks, so you won’t hit those XML parsing errors. You can also easily extract specific elements, text, or attributes from the page using Jsoup’s intuitive API—way more flexible than trying to repurpose XML parsers for HTML.
If You Must Use an XML Parser (Not Recommended)
If you absolutely have to use an XML parser, you’d need to first convert the HTML to well-formed XHTML using a tool like JTidy. But this adds extra complexity and is less reliable than using Jsoup for web content processing.
内容的提问来源于stack exchange,提问作者Teddy

