You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求帮助:编写Java程序抓取网页源码并支持自定义文件名存为TXT

Hey there! Let's get that web scraping and file-saving Java program working for you. I've put together a complete, runnable implementation that checks all your boxes: grabbing webpage source code and letting you save it to a custom-named .txt file.

Working Java Implementation

Here's the full code—just copy this into your project and run it:

import java.io.BufferedWriter;
import java.io.IOException;
import java.io.OutputStreamWriter;
import java.net.URL;
import java.net.URLConnection;
import java.nio.charset.StandardCharsets;
import java.util.Scanner;

public class WebPageScraper {
    public static void main(String[] args) {
        Scanner scanner = new Scanner(System.in);

        // Get target URL from user
        System.out.print("Enter the URL of the webpage you want to scrape: ");
        String urlString = scanner.nextLine().trim();

        // Get custom filename from user
        System.out.print("Enter the custom filename (without .txt extension): ");
        String fileName = scanner.nextLine().trim() + ".txt";

        try {
            // Create URL object and open connection
            URL url = new URL(urlString);
            URLConnection connection = url.openConnection();
            
            // Mimic a browser request to avoid 403 Forbidden errors
            connection.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36");
            
            // Handle character encoding correctly (fallback to UTF-8 if needed)
            String encoding = connection.getContentEncoding();
            if (encoding == null) {
                encoding = StandardCharsets.UTF_8.name();
            }

            // Read the webpage source line by line
            StringBuilder sourceCode = new StringBuilder();
            try (java.io.BufferedReader reader = new java.io.BufferedReader(
                    new java.io.InputStreamReader(connection.getInputStream(), encoding))) {
                String line;
                while ((line = reader.readLine()) != null) {
                    sourceCode.append(line).append(System.lineSeparator());
                }
            }

            // Save the source code to the custom .txt file
            try (BufferedWriter writer = new BufferedWriter(
                    new OutputStreamWriter(new java.io.FileOutputStream(fileName), StandardCharsets.UTF_8))) {
                writer.write(sourceCode.toString());
                System.out.println("Success! Webpage source saved to: " + fileName);
            }

        } catch (java.net.MalformedURLException e) {
            System.err.println("Oops: Invalid URL format. Make sure it starts with http:// or https://!");
        } catch (IOException e) {
            System.err.println("Error: Couldn't read the webpage or save the file. Details: " + e.getMessage());
        } finally {
            scanner.close();
        }
    }
}
Key Features & Fixes

I added these important bits to make sure the program works reliably:

  • Auto-closing streams: Uses try-with-resources to automatically close readers and writers, so you don't have to worry about resource leaks.
  • Browser mimicry: Sets a User-Agent header to avoid getting blocked by websites that reject non-browser requests.
  • Proper encoding handling: Reads the webpage's specified character encoding (or defaults to UTF-8) to ensure the saved text looks exactly like the original source.
  • User-friendly input: Prompts you for both the target URL and custom filename, so you don't have to hardcode values.
Quick Debug Tips

If you still run into issues:

  • Double-check that your URL includes http:// or https://—many sites won't load without this prefix.
  • Make sure your program has permission to write files in the directory where it's running.
  • Verify that the website allows scraping (some sites block bots via robots.txt or other measures).

内容的提问来源于stack exchange,提问作者Yu heng Lim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:52:25