You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从指定URL读取XML遇解析错误,添加UTF-8编码仍未解决

Fixing XML Parsing Error When Trying to Read HTML Content

Hey there, let's break down what's going on here and get your code working properly!

The Core Problem

First off, the root issue is that you’re using an XML parser to try and parse an HTML page. The URL http://www.bnb.bg/ returns regular HTML, not well-formed XML. HTML often has loose syntax (like unquoted attributes, unclosed tags, or elements like <meta> that skip closing tags) that strict XML parsers can’t handle—that’s exactly why you’re seeing the Open quote is expected for attribute "http-equiv" error. Adding UTF-8 encoding won’t fix this, because the problem isn’t character encoding—it’s that the content isn’t valid XML at all.

Bonus Bug in Your Code

On top of that, your code has a NullPointerException waiting to trigger:

Transformer transformer = null;
transformer.setOutputProperty(OutputKeys.OMIT_XML_DECLARATION,"no");

You initialize transformer to null then immediately call a method on it. This line is also unnecessary—you already have the xform Transformer instance, so you should use that instead if you were to continue with XML parsing.

The Solution: Use an HTML Parser

Instead of forcing HTML through an XML parser, use a library built specifically for parsing messy web HTML. Jsoup is the standard choice for Java. Here’s how to rewrite your code to work correctly:

First, add Jsoup to your project. If you’re using Maven, add this dependency to your pom.xml:

<dependency>
    <groupId>org.jsoup</groupId>
    <artifactId>jsoup</artifactId>
    <version>1.17.2</version>
</dependency>

Then, here’s the updated code:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import java.io.IOException;

public class CrawleyCraw {
    public static void main(String[] args) throws IOException {
        String urlString = "http://www.bnb.bg/";
        // Fetch and parse the HTML page with Jsoup (handles messy syntax automatically)
        Document doc = Jsoup.connect(urlString).get();
        
        // Print the pretty-printed HTML
        System.out.println(doc.html());
        
        // Or convert to well-formed XHTML if you need XML-compatible output
        System.out.println(doc.outerHtml());
    }
}

Jsoup takes care of all the unstructured HTML quirks, so you won’t hit those XML parsing errors. You can also easily extract specific elements, text, or attributes from the page using Jsoup’s intuitive API—way more flexible than trying to repurpose XML parsers for HTML.

If you absolutely have to use an XML parser, you’d need to first convert the HTML to well-formed XHTML using a tool like JTidy. But this adds extra complexity and is less reliable than using Jsoup for web content processing.

内容的提问来源于stack exchange,提问作者Teddy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:08:37