You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中校验解析数据语言及确保仅获取英文谷歌新闻数据的方案

Hey there! Let's break down solutions to your two questions using the code and sample data you shared.

1. Validating the Language of Parsed Data

Since the NewsAPI response you provided doesn’t include an explicit language field for each article, you’ll need to detect language from the text content (like title or description) using a reliable language detection library. Here’s a practical approach using the modern, accurate Lingua library:

Step 1: Add the library dependency

If you’re using Maven, add this to your pom.xml:

<dependency>
    <groupId>com.github.pemistahl</groupId>
    <artifactId>lingua</artifactId>
    <version>1.1.0</version> <!-- Use the latest available version -->
</dependency>

Step 2: Modify your code to check language

Update your existing parsing logic to validate each article’s language:

import com.github.pemistahl.lingua.api.Language;
import com.github.pemistahl.lingua.api.LanguageDetector;
import com.github.pemistahl.lingua.api.LanguageDetectorBuilder;

// Initialize detector once (for efficiency) with languages you expect to encounter
LanguageDetector detector = LanguageDetectorBuilder.fromLanguages(
    Language.ENGLISH, Language.RUSSIAN, Language.JAPANESE
).build();

JSONArray rows1 = (JSONArray) rootObject.get("articles");
List<JSONObject> englishArticles = new ArrayList<>();

for (int i = 0; i < rows1.length(); i++) {
    JSONObject article = rows1.getJSONObject(i);
    String title = article.optString("title", "");
    String description = article.optString("description", "");
    
    // Combine title + description for more accurate detection
    String textToAnalyze = title + " " + description;
    
    if (!textToAnalyze.trim().isEmpty()) {
        Language detectedLang = detector.detectLanguageOf(textToAnalyze);
        if (detectedLang == Language.ENGLISH) {
            englishArticles.add(article);
            System.out.println("Keeping English article: " + title);
        } else {
            System.out.println("Skipping non-English article: " + title);
        }
    }
}

Key details:

  • Lingua uses character patterns and word frequency to minimize false detections.
  • Combining title and description avoids issues if one field is empty or too short.
2. Ensuring You Only Fetch English Data Directly from the API

The most efficient way to get English-only news is to filter at the API level—this cuts down on data transfer and parsing work. NewsAPI supports a language parameter for exactly this purpose.

Step 1: Update your API request URL

Add &language=en to your existing URL (using the ISO 639-1 code for English):

String url = "https://newsapi.org/v2/top-headlines?sources=google-news&apiKey=89c8009165774e0fad3742f78b50c6da&language=en";

Step 2: Add validation as a fallback

Even with the language parameter, it’s smart to keep the language detection logic from Section 1 as a safety net. This handles edge cases where the API might return misclassified content.

Note on the language parameter:

  • It filters results to only include articles published in the specified language.
  • Works with all NewsAPI endpoints that return article content (like top-headlines).

内容的提问来源于stack exchange,提问作者Tanu Garg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:56:36