网站PHP Google News阅读器问题:图片不显示及仅保留主标题首源文章
Fixing PHP Google News Reader Issues: Missing Images & Duplicate Title Filtering
Hey there! Let’s work through both of your problems with straightforward, actionable fixes:
1. Troubleshooting Missing Images
Images failing to load typically boils down to one of these common issues:
- Relative URL Paths: If your scraper pulls partial image URLs (like
/images/xyz.jpginstead of the fullhttps://news.google.com/images/xyz.jpg), your browser can’t resolve them. Double-check that your code prepends the Google News base URL to any image sources you extract. - Hotlink Protection: Google blocks direct embedding of their images from external sites. To bypass this, proxy images through your server using PHP:
Use this function to output images as base64 strings in yourfunction proxy_google_image($image_url) { $curl = curl_init($image_url); curl_setopt($curl, CURLOPT_RETURNTRANSFER, true); curl_setopt($curl, CURLOPT_FOLLOWLOCATION, true); $image_data = curl_exec($curl); curl_close($curl); return 'data:image/jpeg;base64,' . base64_encode($image_data); }<img>tags (e.g.,<img src="<?php echo proxy_google_image($img_url); ?>">). - Incorrect Selector: Ensure your scraping logic targets the right HTML element for images. Google News uses specific classes for article thumbnails—verify your selector (like
$article->find('img.tvs3Id', 0)if using Simple HTML DOM).
2. Filtering to Keep Only the First Article per Main Title
To retain just the first source article for each unique headline, track titles as you process results:
Here’s a practical implementation:
$seen_titles = []; $unique_articles = []; // Assume $raw_articles is your array of fetched Google News entries foreach ($raw_articles as $article) { $main_headline = trim($article['main_title']); // Adjust to match your title extraction method // Skip if we've already added this headline if (!in_array($main_headline, $seen_titles)) { $seen_titles[] = $main_headline; $unique_articles[] = $article; // Stop once we have 5 unique articles (your desired limit) if (count($unique_articles) === 5) { break; } } } // Use $unique_articles to render your final list
This code creates a tracker for seen titles, skips duplicates, and stops once you hit your 5-article limit. Just make sure $article['main_title'] targets the core headline (not source-specific subheads).
内容的提问来源于stack exchange,提问作者nickko
相关产品推荐
相关产品推荐

