You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Java中获取指定站点资源目录下的静态文件列表?

Great question! Let's break this down clearly since there's a key caveat you need to know first:

获取第三方站点目录下静态文件列表的Java方案

First off, a critical reality: most public web servers (including Naver's static server) don't enable directory indexing. If you try to directly request https://static.naver.com/images/, you'll almost certainly get a 403 Forbidden or 404 error instead of a file list. So we need to work around this with alternative approaches.

Practical Solutions

If those images are referenced on any Naver webpage, you can crawl that page's HTML, parse it, and extract links pointing to the https://static.naver.com/images/ directory.

Here's an example using the Jsoup library for HTML parsing:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;
import java.io.IOException;
import java.util.HashSet;
import java.util.Set;

public class ImageLinkScraper {
    public static void main(String[] args) {
        String targetPage = "https://www.naver.com"; // Assume images are referenced here
        String imageBasePath = "https://static.naver.com/images/";
        Set<String> collectedImageUrls = new HashSet<>();

        try {
            // Fetch and parse the target page
            Document pageDoc = Jsoup.connect(targetPage).get();
            // Grab all img tags with a src attribute
            Elements imgTags = pageDoc.select("img[src]");
            
            for (Element img : imgTags) {
                String fullImgUrl = img.attr("abs:src"); // Get absolute URL instead of relative
                if (fullImgUrl.startsWith(imageBasePath)) {
                    collectedImageUrls.add(fullImgUrl);
                }
            }

            // Output the results
            System.out.println("Found images under " + imageBasePath + ":");
            for (String url : collectedImageUrls) {
                System.out.println("- " + url);
            }

        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

To use Jsoup, add this Maven dependency to your project:

<dependency>
    <groupId>org.jsoup</groupId>
    <artifactId>jsoup</artifactId>
    <version>1.17.2</version>
</dependency>

2. Check for public resource manifests

Some sites publish public resource lists (like manifest.json, sitemap.xml, or dedicated API endpoints) that list static assets. You can try accessing these files to extract image paths:

  • https://static.naver.com/images/manifest.json
  • https://static.naver.com/sitemap.xml

If these exist, use Java's built-in HttpClient to fetch and parse the content, then pull out the image URLs.

If you know the naming pattern of the images (like your example's pencil.png, note.png), you can construct URLs and send HEAD requests to check if they exist:

import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.Arrays;
import java.util.List;

public class ImageExistenceChecker {
    public static void main(String[] args) {
        String baseUrl = "https://static.naver.com/images/";
        List<String> possibleImageNames = Arrays.asList("pencil.png", "note.png", "table.png", "book.png");

        HttpClient client = HttpClient.newHttpClient();

        for (String name : possibleImageNames) {
            String fullUrl = baseUrl + name;
            HttpRequest request = HttpRequest.newBuilder()
                    .uri(URI.create(fullUrl))
                    .method("HEAD", HttpRequest.BodyPublishers.noBody())
                    .build();

            try {
                HttpResponse<Void> response = client.send(request, HttpResponse.BodyHandlers.discarding());
                if (response.statusCode() == 200) {
                    System.out.println("Image exists: " + fullUrl);
                }
            } catch (Exception e) {
                e.printStackTrace();
            }
        }
    }
}

⚠️ Warning: This method can trigger anti-scraping measures on the target server, leading to your IP being blocked. It's inefficient and not recommended for large-scale use.

Key Reminders

  • Always check the target site's robots.txt (e.g., https://static.naver.com/robots.txt) before scraping to avoid violating their terms of service or legal rules.
  • If the server doesn't provide a public file list, you can't get a complete list of all images in the directory—you'll only be able to collect images that are referenced publicly or match your guessed naming patterns.

内容的提问来源于stack exchange,提问作者jini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:01:40