You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为路透社及滚动加载型链接设置日期过滤器,仅抓取近1年的新闻文章

Date Filtering Strategies for Two News Scraping Tasks

1. Scraping Reuters News After a Specific Date

Option 1: Use Reuters' Built-in Date Filters (Most Efficient)

  • First, navigate to Reuters' search page or the specific section you want to scrape. Look for date range filters (usually labeled "Date" or "Timeframe").
  • When you set your target start date, check the URL—Reuters often encodes date parameters like afterDate=YYYYMMDD or startDate=YYYY-MM-DD in the query string. You can directly use this URL in your scraper to only fetch results after your chosen date, skipping post-scraping filtering entirely.
  • Example URL structure (hypothetical): https://www.reuters.com/search/news?blob=your_topic&afterDate=20240101

Option 2: Post-Scraping Date Filtering (If No Built-in Filter)

If the section you're scraping lacks a native date filter, extract and validate each article's publish date manually:

  • Step 1: Pull the publish date from each article. Reuters typically uses formats like Jan 1, 2024 or 2024-01-01—use a parser (like Python's datetime module) to convert this string into a comparable date object.
  • Step 2: Compare the parsed date against your target cutoff. Discard any articles where the publish date falls before your specified date.

Quick Python snippet (using BeautifulSoup):

from datetime import datetime
from bs4 import BeautifulSoup
import requests

target_date = datetime(2024, 1, 1)  # Your custom cutoff date
url = "https://www.reuters.com/your-target-section"

response = requests.get(url, headers={"User-Agent": "Mozilla/5.0"})
soup = BeautifulSoup(response.text, "html.parser")

for article in soup.find_all("article"):
    date_str = article.find("time")["datetime"]
    publish_date = datetime.fromisoformat(date_str.replace("Z", "+00:00"))
    if publish_date >= target_date:
        # Process the valid article content
        print(f"Valid article: {article.find('h3').text.strip()}")
    else:
        # Skip outdated articles, or stop if paginating
        continue

2. Scraping Infinite-Scroll Pages for the Last 12 Months of News

Infinite-scroll pages load older content as you scroll, so your filter needs to halt scraping once you hit articles older than 1 year. Here's how to implement this:

Step 1: Calculate Your 1-Year Cutoff Date

First, define the cutoff relative to today:

from datetime import datetime, timedelta

one_year_ago = datetime.now() - timedelta(days=365)

Step 2: Handle Infinite Scroll & Date Checks

If the Page Uses API-Based Loading (Best Case)

Many infinite-scroll sites load content via an API endpoint (check the Network tab in your browser's DevTools). Look for parameters like before (Unix timestamp) or max_date in API requests. You can:

  • Start with the current date to fetch batches of articles.
  • After each batch, check the oldest article's date. If it's earlier than one_year_ago, stop requesting more batches.
  • Example API request (hypothetical): https://example.com/api/news?limit=20&before=1704067200 (Unix timestamp for Jan 1, 2024)

If the Page Uses Frontend-Rendered Scrolling (Use Browser Automation)

If content is rendered client-side (no exposed API), use tools like Selenium or Playwright to simulate scrolling:

  • Scroll incrementally, wait for new content to load.
  • After each scroll, extract all visible articles' dates.
  • Stop scrolling as soon as any article's date is earlier than one_year_ago (checking the oldest date in the current batch is more efficient).

Python snippet with Selenium:

from selenium import webdriver
from selenium.webdriver.common.by import By
from datetime import datetime, timedelta
import time

one_year_ago = datetime.now() - timedelta(days=365)
driver = webdriver.Chrome()
driver.get("https://example.com/infinite-scroll-news")

seen_articles = set()  # Avoid duplicate entries

while True:
    # Extract current visible articles
    articles = driver.find_elements(By.TAG_NAME, "article")
    for article in articles:
        article_link = article.find_element(By.TAG_NAME, "a").get_attribute("href")
        if article_link in seen_articles:
            continue
        seen_articles.add(article_link)
        
        date_str = article.find_element(By.TAG_NAME, "time").get_attribute("datetime")
        publish_date = datetime.fromisoformat(date_str.replace("Z", "+00:00"))
        
        if publish_date < one_year_ago:
            # Stop scraping immediately
            driver.quit()
            print("Reached articles older than 1 year—stopping process.")
            exit()
        else:
            # Process the valid article
            print(f"Processing: {article.find_element(By.TAG_NAME, 'h2').text.strip()}")
    
    # Scroll to load more content
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # Adjust wait time based on page load speed

Key Notes for Both Tasks

  • Timezone Consistency: Always normalize dates to the same timezone (e.g., UTC) to avoid incorrect comparisons. Reuters uses UTC in most datetime attributes, so align your cutoff date to UTC if needed.
  • Anti-Scraping Measures: Add delays between requests, use a valid User-Agent header, and consider rotating proxies if you hit rate limits.
  • Date Parsing Edge Cases: Some articles might use relative dates (e.g., "2 hours ago")—convert these to absolute dates using libraries like dateutil.parser.

内容的提问来源于stack exchange,提问作者huy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 12:27:32