InputStreamReader与URL到底是什么?附Java网页读取代码
关于URL类与InputStreamReader的作用疑问
我这段代码能正常运行,但它是参考网页读取教程写的。我之前只粗浅接触过File I/O,从没碰到过这种网络读取的场景,实在搞不懂InputStreamReader和URL标识符的作用。
import java.net.*; import java.util.*; import java.io.*; public class urlReader { public static void main(String[] args) { URL[] websites = new URL[4]; URLConnection conn = null; // Try catch statement that handles no URL found and i/o exceptions // try { // a list of websites sorted into arrays // websites[0] = new URL("https://www.pravdareport.com"); websites[1] = new URL("https://pravda.ru"); websites[2] = new URL("https://www.lefigaro.fr"); websites[3] = new URL("https://www.independent.co.uk"); // For each website it assigns the websites to webpage after each iteration of the entire loop // for (int i = 0; i < websites.length; i++) { URL webpage = websites[i]; conn = webpage.openConnection(); // This is used so that for each website it ''opens'' the connection // InputStreamReader reader = new InputStreamReader(conn.getInputStream(), "UTF8"); // BufferedReader br = new BufferedReader(reader); String lines = ""; while ((lines = br.readLine()) != null) { if (lines.indexOf("<title>") != -1 && lines.indexOf("</title>") != -1) { String title1 = lines.substring(lines.indexOf("<title>") + 7, lines.indexOf("</title>")); System.out.println("The Title of pravdareport.com is: " + title1); break; } } } } catch (MalformedURLException e) { e.printStackTrace(); } catch (IOException e) { e.printStackTrace(); } } }
1. URL类的作用
URL是Java用来封装网络资源地址的核心类,类比你本地文件的路径——本地路径指向硬盘上的文件,URL对象则指向互联网上的一个资源(比如网页、接口、图片):
- 你传入的网址字符串会被自动解析成标准格式,Java帮你处理协议(http/https)、域名、端口这些底层细节,不用你手动拆分校验。
- 调用
openConnection()时,URL会帮你完成与目标服务器的连接建立(类似浏览器输入网址后的握手过程),返回的URLConnection就是这个连接的抽象,通过它你能拿到服务器返回的所有数据。
2. InputStreamReader的作用
你之前接触的File I/O读的是本地文件,而网络请求拿到的是原始字节流(InputStream),但我们需要的是可读的文本内容,InputStreamReader就是干这个转换的:
- 它是字节流到字符流的转换器:服务器返回的网页数据是UTF-8编码的字节,
InputStreamReader会按照你指定的编码(这里是UTF8)把字节转成Java能识别的字符。 - 搭配
BufferedReader是为了提升效率:InputStreamReader是逐个字符读取,速度慢;BufferedReader会先缓存一批数据,再按行读取,这也是你能用br.readLine()按行扫描网页内容的原因。
小建议
你代码里打印标题时固定写了pravdareport.com,可以改成动态显示当前网站的域名,更通用:
System.out.println("The Title of " + webpage.getHost() + " is: " + title1);
内容的提问来源于stack exchange,提问作者userr554
相关产品推荐
相关产品推荐

