You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Jsoup遍历多URL时遇非法字符错误,求解决方法

Fixing "Illegal character in the URL" Error for Your Web Crawler

Hey Jeff, let's break down why you're hitting that frustrating "Illegal character in the URL" error and get your crawler back up and running. This issue almost always ties back to how you're reading or processing the URLs from your text file—here are the most common fixes to try:

1. Fix Character Encoding When Reading the File

The readAllBytes method returns raw byte data, and converting it to a String without specifying an encoding uses your system's default charset (which might not match how your text file was saved). This can mangle special characters like accented letters, spaces, or URL-encoded symbols.

Fixed code for reading the file:

import java.nio.file.Files;
import java.nio.file.Paths;
import java.nio.charset.StandardCharsets;

String fileContent = new String(
    Files.readAllBytes(Paths.get("your-urls-file.txt")),
    StandardCharsets.UTF_8 // Use the same encoding your text file was saved with (UTF-8 is safest)
);

2. Clean Up Hidden/Whitespace Characters in URLs

Text files often have invisible characters like newlines (\n), carriage returns (\r), or leading/trailing spaces that sneak into your URL strings. These are the #1 culprit for "illegal character" errors.

Code to sanitize URLs after reading:

import java.util.Arrays;
import java.util.List;
import java.util.stream.Collectors;

// Split content into URLs (handles both Windows and Linux line endings)
List<String> cleanUrls = Arrays.stream(fileContent.split("\\r?\\n"))
    .map(String::trim) // Remove leading/trailing spaces, tabs, etc.
    .filter(url -> !url.isEmpty()) // Skip empty lines
    .collect(Collectors.toList());

3. Properly Escape Invalid URL Characters

If your URLs contain characters like spaces, &, ?, or non-ASCII symbols, they need to be URL-encoded. Important: don't encode the entire URL (this will break the http:// part)—only encode the path or query parameter sections.

Helper method to sanitize individual URLs:

import java.net.URL;
import java.net.URLEncoder;
import java.io.UnsupportedEncodingException;

public String sanitizeSingleUrl(String rawUrl) throws UnsupportedEncodingException {
    URL url = new URL(rawUrl);
    
    // Encode the path, but preserve slashes and fix space encoding
    String encodedPath = URLEncoder.encode(url.getPath(), "UTF-8")
        .replace("+", "%20") // URLEncoder uses "+" for spaces; URLs require "%20"
        .replace("%2F", "/"); // Don't encode the path separator "/"
    
    // Reconstruct the clean URL (include query params if present)
    return new URL(
        url.getProtocol(),
        url.getHost(),
        url.getPort(),
        encodedPath + (url.getQuery() != null ? "?" + url.getQuery() : "")
    ).toString();
}

Debug Tip to Pinpoint the Culprit

Before passing URLs to your crawler, print them wrapped in quotes to spot hidden characters:

for (String url : cleanUrls) {
    System.out.println("Processing URL: '" + url + "'");
}

Look for things like 'http://example.com/page 1' (unencoded space) or 'http://example.com/page\n' (hidden newline)—these are the illegal characters triggering the error.

内容的提问来源于stack exchange,提问作者Jeff

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:39:13