You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java程序下载XKCD前100幅漫画问题求助

Fixing Your XKCD Comic Downloader in Java

Hey there! Let's get your program downloading those first 100 XKCD comics properly. I'll walk you through the key issues and a working solution, covering both a line-based approach (since you mentioned the img tag is on line 50) and a more robust HTML parsing method that avoids relying on line numbers.

Common Pitfalls You Might Be Hitting

  • Relative URL Handling: The img src in XKCD's page is a relative path (like //imgs.xkcd.com/comics/foo.png), so you need to convert it to a full HTTPS URL to download the image.
  • HTTPS Connection Setup: Java requires proper SSL handling for HTTPS requests, though modern Java versions make this seamless for well-known sites like XKCD.
  • Line Counting Inaccuracy: Hidden newlines or whitespace in the HTML might throw off your line number count—line 50 might not be what you expect if you're not reading the raw source correctly.
  • Unclosed Resources: Forgetting to close input/output streams can cause silent failures or resource leaks.

Working Solution: Line-Based Approach (Matching Your Requirement)

This code reads each comic page, targets line 50, extracts the image URL, and saves the comic to your local directory:

import java.io.*;
import java.net.*;
import java.util.regex.Matcher;
import java.util.regex.Pattern;

public class XKCDDownloader {
    // Regex to match the image source in the HTML line
    private static final Pattern IMG_SRC_PATTERN = Pattern.compile("src=\"(//imgs\\.xkcd\\.com/comics/.*?)\"");

    public static void main(String[] args) {
        // Update this to your desired save directory
        String saveDirectory = "./xkcd_comics/";
        new File(saveDirectory).mkdirs(); // Create directory if it doesn't exist

        for (int comicNumber = 1; comicNumber <= 100; comicNumber++) {
            try {
                // Step 1: Fetch the comic page HTML
                URL pageUrl = new URL("https://xkcd.com/" + comicNumber + "/");
                HttpURLConnection connection = (HttpURLConnection) pageUrl.openConnection();
                connection.setRequestMethod("GET");
                connection.setInstanceFollowRedirects(true);

                BufferedReader htmlReader = new BufferedReader(
                        new InputStreamReader(connection.getInputStream())
                );

                String line;
                int lineCount = 0;
                String imageSrc = null;

                // Step 2: Find line 50 and extract the image URL
                while ((line = htmlReader.readLine()) != null) {
                    lineCount++;
                    if (lineCount == 50) {
                        Matcher matcher = IMG_SRC_PATTERN.matcher(line);
                        if (matcher.find()) {
                            // Convert relative URL to full HTTPS URL
                            imageSrc = "https:" + matcher.group(1);
                        }
                        break; // Stop reading once we hit line 50
                    }
                }

                htmlReader.close();
                connection.disconnect();

                if (imageSrc == null) {
                    System.out.println("Failed to find image URL for comic #" + comicNumber);
                    continue;
                }

                // Step 3: Download and save the image
                URL imageUrl = new URL(imageSrc);
                InputStream imageStream = imageUrl.openStream();
                FileOutputStream outputStream = new FileOutputStream(
                        saveDirectory + "xkcd_" + comicNumber + ".png"
                );

                byte[] buffer = new byte[4096];
                int bytesRead;
                while ((bytesRead = imageStream.read(buffer)) != -1) {
                    outputStream.write(buffer, 0, bytesRead);
                }

                imageStream.close();
                outputStream.close();
                System.out.println("Successfully downloaded comic #" + comicNumber);

            } catch (IOException e) {
                System.out.println("Error downloading comic #" + comicNumber + ": " + e.getMessage());
                e.printStackTrace();
            }
        }
    }
}

Key Details:

  • Regex Matching: The regex targets the exact format of XKCD's image paths, ensuring we grab the correct src value without extra noise.
  • Relative to Absolute URL: We prepend https: to the relative src to form a valid, downloadable URL.
  • Auto-Directory Creation: The code creates your save folder automatically if it doesn't exist.
  • Basic Error Handling: Exception catching prevents the program from crashing if one comic fails to download.

More Robust Solution: Use Jsoup (No Line Dependency)

Relying on line numbers is fragile—if XKCD updates their page structure, line 50 won't hold the image tag anymore. Using Jsoup (a lightweight Java HTML parser) makes this far more reliable:

  1. Add Jsoup to your project (via Maven or by downloading the JAR).
  2. Replace the line-reading logic with this snippet:
// Fetch and parse the page with Jsoup
org.jsoup.nodes.Document doc = org.jsoup.Jsoup.connect("https://xkcd.com/" + comicNumber + "/").get();
org.jsoup.nodes.Element imgElement = doc.select("#comic img").first();

if (imgElement == null) {
    System.out.println("Failed to find image for comic #" + comicNumber);
    continue;
}
// Get the full HTTPS URL automatically
String imageSrc = imgElement.absUrl("src");

This targets the image directly using its CSS selector (#comic img), which is stable even if the page's line structure changes.

Final Tips

  • Ensure you're using Java 8 or newer for smooth HTTPS support.
  • If you run into SSL certificate errors (unlikely for XKCD), you may need to configure an SSLContext, but this is rare for mainstream sites.
  • Always close streams and connections to avoid resource leaks.

内容的提问来源于stack exchange,提问作者Trevor Richerson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:39:57