Java中校验解析数据语言及确保仅获取英文谷歌新闻数据的方案
Hey there! Let's break down solutions to your two questions using the code and sample data you shared.
Since the NewsAPI response you provided doesn’t include an explicit language field for each article, you’ll need to detect language from the text content (like title or description) using a reliable language detection library. Here’s a practical approach using the modern, accurate Lingua library:
Step 1: Add the library dependency
If you’re using Maven, add this to your pom.xml:
<dependency> <groupId>com.github.pemistahl</groupId> <artifactId>lingua</artifactId> <version>1.1.0</version> <!-- Use the latest available version --> </dependency>
Step 2: Modify your code to check language
Update your existing parsing logic to validate each article’s language:
import com.github.pemistahl.lingua.api.Language; import com.github.pemistahl.lingua.api.LanguageDetector; import com.github.pemistahl.lingua.api.LanguageDetectorBuilder; // Initialize detector once (for efficiency) with languages you expect to encounter LanguageDetector detector = LanguageDetectorBuilder.fromLanguages( Language.ENGLISH, Language.RUSSIAN, Language.JAPANESE ).build(); JSONArray rows1 = (JSONArray) rootObject.get("articles"); List<JSONObject> englishArticles = new ArrayList<>(); for (int i = 0; i < rows1.length(); i++) { JSONObject article = rows1.getJSONObject(i); String title = article.optString("title", ""); String description = article.optString("description", ""); // Combine title + description for more accurate detection String textToAnalyze = title + " " + description; if (!textToAnalyze.trim().isEmpty()) { Language detectedLang = detector.detectLanguageOf(textToAnalyze); if (detectedLang == Language.ENGLISH) { englishArticles.add(article); System.out.println("Keeping English article: " + title); } else { System.out.println("Skipping non-English article: " + title); } } }
Key details:
- Lingua uses character patterns and word frequency to minimize false detections.
- Combining title and description avoids issues if one field is empty or too short.
The most efficient way to get English-only news is to filter at the API level—this cuts down on data transfer and parsing work. NewsAPI supports a language parameter for exactly this purpose.
Step 1: Update your API request URL
Add &language=en to your existing URL (using the ISO 639-1 code for English):
String url = "https://newsapi.org/v2/top-headlines?sources=google-news&apiKey=89c8009165774e0fad3742f78b50c6da&language=en";
Step 2: Add validation as a fallback
Even with the language parameter, it’s smart to keep the language detection logic from Section 1 as a safety net. This handles edge cases where the API might return misclassified content.
Note on the language parameter:
- It filters results to only include articles published in the specified language.
- Works with all NewsAPI endpoints that return article content (like
top-headlines).
内容的提问来源于stack exchange,提问作者Tanu Garg

