You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python newspaper包返回文章的判定规则及CNN网址异常咨询

Hey there! Let's tackle your questions about the Python newspaper library clearly and directly.

1. What articles does the Python newspaper package return?

The newspaper library is built to fetch journalistic-style articles—pages that have core news attributes like a clear headline, substantial body text, publication dates, author information, and are designed to deliver news content. It automatically filters out non-article pages like navigation menus, ad landing pages, static informational pages, or empty/short content pages that don't fit the "news article" profile.

2. How does newspaper decide which URLs/articles to return? (And why CNN Politics gives the same results as the main site?)

This is a common gotcha with large news sites like CNN, and it boils down to how newspaper prioritizes content and analyzes site structures. Here's the breakdown of its core logic:

  • Site-wide feed prioritization: When you pass a URL (even a sub-section like https://www.cnn.com/politics), newspaper first tries to detect and pull from the site's RSS feeds or sitemaps. Many big news sites (including CNN) feature their most popular/latest articles across multiple feeds—so the politics section's top stories often appear in the main site's feed too. If newspaper defaults to grabbing this global popular feed instead of the politics-specific feed, you'll get the same articles as the main site.
  • Article feature detection: The library uses algorithms to judge whether a page qualifies as a news article. It looks for things like word count thresholds, presence of publication metadata, and structural cues (like article-specific HTML tags). It won't waste time crawling non-article pages, but if it doesn't recognize a sub-section's list page as a source of article links, it won't dig deeper into that section.
  • Default limited crawl depth: newspaper doesn't crawl every single page under a sub-section by default. It's optimized to get quick, relevant results (usually the latest/hottest articles) rather than performing a full site scrape. So when you target cnn.com/politics, it might not traverse every article link on that section's pages—instead, it falls back to the site's primary feed.

Fixes to get section-specific articles

If you want to strictly pull articles from the CNN Politics section, try these tweaks:

  • Use the section's dedicated RSS feed: Find the official RSS URL for CNN Politics (most news sections have one) and pass that to newspaper.build() instead of the web page URL. This ensures you're only getting content from that section.
  • Adjust crawl parameters: Disable caching with memoize_articles=False to avoid reusing old results, and increase the crawl depth if needed to let the library traverse more links in the sub-section.
  • Filter links manually: After fetching initial links, add a check to only keep URLs that include /politics/ in their path before parsing them into articles.

内容的提问来源于stack exchange,提问作者r1234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:22:59