如何通过Java调用Stanford CoreNLP Server获取处理结果
解决HttpUrlConnection访问Stanford CoreNLP Server返回FileNotFoundException的问题
首先,你遇到的FileNotFoundException并不是真的找不到文件,而是HttpUrlConnection的默认行为:当服务器返回非2xx的状态码时,直接调用getInputStream()就会抛出这个异常。我们需要先检查响应状态码,再处理输入流;同时还要确保Java请求的格式和参数和你的wget命令完全对齐。
完整的Java实现代码
import java.io.BufferedReader; import java.io.IOException; import java.io.InputStreamReader; import java.io.OutputStream; import java.net.HttpURLConnection; import java.net.URL; import java.net.URLEncoder; import java.nio.charset.StandardCharsets; import java.nio.file.Files; import java.nio.file.Paths; public class CoreNLPClient { public static void main(String[] args) { String text = "It is a nice day, isn't it?"; String serverUrl = "http://192.168.1.30:9000/"; String properties = "{\"annotators\": \"openie\", \"outputFormat\": \"json\"}"; try { // 对properties的JSON字符串做URL编码,避免特殊字符导致请求错误 String encodedProperties = URLEncoder.encode(properties, StandardCharsets.UTF_8.name()); String fullUrl = serverUrl + "?properties=" + encodedProperties; URL url = new URL(fullUrl); HttpURLConnection conn = (HttpURLConnection) url.openConnection(); // 设置POST请求方法 conn.setRequestMethod("POST"); // 允许写入请求体 conn.setDoOutput(true); // 设置Content-Type,和wget默认的POST格式一致 conn.setRequestProperty("Content-Type", "application/x-www-form-urlencoded; charset=UTF-8"); // 将待分析的文本写入请求体 try (OutputStream os = conn.getOutputStream()) { byte[] inputBytes = text.getBytes(StandardCharsets.UTF_8); os.write(inputBytes, 0, inputBytes.length); } // 先获取响应状态码,判断请求是否成功 int responseCode = conn.getResponseCode(); System.out.println("服务器响应码: " + responseCode); // 根据响应码选择对应的输入流(成功用正常流,失败用错误流) BufferedReader reader; if (responseCode >= 200 && responseCode < 300) { reader = new BufferedReader(new InputStreamReader(conn.getInputStream(), StandardCharsets.UTF_8)); } else { reader = new BufferedReader(new InputStreamReader(conn.getErrorStream(), StandardCharsets.UTF_8)); } // 读取响应内容 StringBuilder responseContent = new StringBuilder(); String line; while ((line = reader.readLine()) != null) { responseContent.append(line); } reader.close(); // 将结果写入res.txt文件,和wget的效果一致 Files.write(Paths.get("res.txt"), responseContent.toString().getBytes(StandardCharsets.UTF_8)); System.out.println("结果已保存至res.txt"); } catch (IOException e) { e.printStackTrace(); } } }
核心注意事项
- 处理响应状态码:必须先调用
conn.getResponseCode(),再根据状态码选择输入流,避免直接抛出异常。如果请求失败,getErrorStream()能帮你看到服务器返回的具体错误原因。 - URL编码参数:properties里的JSON包含引号等特殊字符,必须用
URLEncoder.encode()编码后再拼入URL,否则服务器会解析错误。 - 对齐请求格式:设置和wget一致的
Content-Type,确保服务器能正确识别POST体中的文本内容。 - 编码一致性:全程使用UTF-8编码,避免出现乱码问题。
排查小技巧
如果还是有问题,可以:
- 打印出拼接后的完整URL,和wget命令里的URL对比,确保参数完全一致。
- 查看
getErrorStream()返回的内容,服务器通常会明确告诉你哪里出错了(比如参数格式错误、端口不可达等)。 - 用curl命令测试请求,对比Java请求的差异:
curl --data 'It is a nice day, isn'\''t it?' 'http://192.168.1.30:9000/?properties={"annotators": "openie", "outputFormat": "json"}'
内容的提问来源于stack exchange,提问作者Hey
相关产品推荐
相关产品推荐

