求帮助:编写Java程序抓取网页源码并支持自定义文件名存为TXT
Hey there! Let's get that web scraping and file-saving Java program working for you. I've put together a complete, runnable implementation that checks all your boxes: grabbing webpage source code and letting you save it to a custom-named .txt file.
Working Java Implementation
Here's the full code—just copy this into your project and run it:
import java.io.BufferedWriter; import java.io.IOException; import java.io.OutputStreamWriter; import java.net.URL; import java.net.URLConnection; import java.nio.charset.StandardCharsets; import java.util.Scanner; public class WebPageScraper { public static void main(String[] args) { Scanner scanner = new Scanner(System.in); // Get target URL from user System.out.print("Enter the URL of the webpage you want to scrape: "); String urlString = scanner.nextLine().trim(); // Get custom filename from user System.out.print("Enter the custom filename (without .txt extension): "); String fileName = scanner.nextLine().trim() + ".txt"; try { // Create URL object and open connection URL url = new URL(urlString); URLConnection connection = url.openConnection(); // Mimic a browser request to avoid 403 Forbidden errors connection.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"); // Handle character encoding correctly (fallback to UTF-8 if needed) String encoding = connection.getContentEncoding(); if (encoding == null) { encoding = StandardCharsets.UTF_8.name(); } // Read the webpage source line by line StringBuilder sourceCode = new StringBuilder(); try (java.io.BufferedReader reader = new java.io.BufferedReader( new java.io.InputStreamReader(connection.getInputStream(), encoding))) { String line; while ((line = reader.readLine()) != null) { sourceCode.append(line).append(System.lineSeparator()); } } // Save the source code to the custom .txt file try (BufferedWriter writer = new BufferedWriter( new OutputStreamWriter(new java.io.FileOutputStream(fileName), StandardCharsets.UTF_8))) { writer.write(sourceCode.toString()); System.out.println("Success! Webpage source saved to: " + fileName); } } catch (java.net.MalformedURLException e) { System.err.println("Oops: Invalid URL format. Make sure it starts with http:// or https://!"); } catch (IOException e) { System.err.println("Error: Couldn't read the webpage or save the file. Details: " + e.getMessage()); } finally { scanner.close(); } } }
Key Features & Fixes
I added these important bits to make sure the program works reliably:
- Auto-closing streams: Uses
try-with-resourcesto automatically close readers and writers, so you don't have to worry about resource leaks. - Browser mimicry: Sets a
User-Agentheader to avoid getting blocked by websites that reject non-browser requests. - Proper encoding handling: Reads the webpage's specified character encoding (or defaults to UTF-8) to ensure the saved text looks exactly like the original source.
- User-friendly input: Prompts you for both the target URL and custom filename, so you don't have to hardcode values.
Quick Debug Tips
If you still run into issues:
- Double-check that your URL includes
http://orhttps://—many sites won't load without this prefix. - Make sure your program has permission to write files in the directory where it's running.
- Verify that the website allows scraping (some sites block bots via
robots.txtor other measures).
内容的提问来源于stack exchange,提问作者Yu heng Lim
相关产品推荐
相关产品推荐

