如何为路透社及滚动加载型链接设置日期过滤器,仅抓取近1年的新闻文章
1. Scraping Reuters News After a Specific Date
Option 1: Use Reuters' Built-in Date Filters (Most Efficient)
- First, navigate to Reuters' search page or the specific section you want to scrape. Look for date range filters (usually labeled "Date" or "Timeframe").
- When you set your target start date, check the URL—Reuters often encodes date parameters like
afterDate=YYYYMMDDorstartDate=YYYY-MM-DDin the query string. You can directly use this URL in your scraper to only fetch results after your chosen date, skipping post-scraping filtering entirely. - Example URL structure (hypothetical):
https://www.reuters.com/search/news?blob=your_topic&afterDate=20240101
Option 2: Post-Scraping Date Filtering (If No Built-in Filter)
If the section you're scraping lacks a native date filter, extract and validate each article's publish date manually:
- Step 1: Pull the publish date from each article. Reuters typically uses formats like
Jan 1, 2024or2024-01-01—use a parser (like Python'sdatetimemodule) to convert this string into a comparable date object. - Step 2: Compare the parsed date against your target cutoff. Discard any articles where the publish date falls before your specified date.
Quick Python snippet (using BeautifulSoup):
from datetime import datetime from bs4 import BeautifulSoup import requests target_date = datetime(2024, 1, 1) # Your custom cutoff date url = "https://www.reuters.com/your-target-section" response = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}) soup = BeautifulSoup(response.text, "html.parser") for article in soup.find_all("article"): date_str = article.find("time")["datetime"] publish_date = datetime.fromisoformat(date_str.replace("Z", "+00:00")) if publish_date >= target_date: # Process the valid article content print(f"Valid article: {article.find('h3').text.strip()}") else: # Skip outdated articles, or stop if paginating continue
2. Scraping Infinite-Scroll Pages for the Last 12 Months of News
Infinite-scroll pages load older content as you scroll, so your filter needs to halt scraping once you hit articles older than 1 year. Here's how to implement this:
Step 1: Calculate Your 1-Year Cutoff Date
First, define the cutoff relative to today:
from datetime import datetime, timedelta one_year_ago = datetime.now() - timedelta(days=365)
Step 2: Handle Infinite Scroll & Date Checks
If the Page Uses API-Based Loading (Best Case)
Many infinite-scroll sites load content via an API endpoint (check the Network tab in your browser's DevTools). Look for parameters like before (Unix timestamp) or max_date in API requests. You can:
- Start with the current date to fetch batches of articles.
- After each batch, check the oldest article's date. If it's earlier than
one_year_ago, stop requesting more batches. - Example API request (hypothetical):
https://example.com/api/news?limit=20&before=1704067200(Unix timestamp for Jan 1, 2024)
If the Page Uses Frontend-Rendered Scrolling (Use Browser Automation)
If content is rendered client-side (no exposed API), use tools like Selenium or Playwright to simulate scrolling:
- Scroll incrementally, wait for new content to load.
- After each scroll, extract all visible articles' dates.
- Stop scrolling as soon as any article's date is earlier than
one_year_ago(checking the oldest date in the current batch is more efficient).
Python snippet with Selenium:
from selenium import webdriver from selenium.webdriver.common.by import By from datetime import datetime, timedelta import time one_year_ago = datetime.now() - timedelta(days=365) driver = webdriver.Chrome() driver.get("https://example.com/infinite-scroll-news") seen_articles = set() # Avoid duplicate entries while True: # Extract current visible articles articles = driver.find_elements(By.TAG_NAME, "article") for article in articles: article_link = article.find_element(By.TAG_NAME, "a").get_attribute("href") if article_link in seen_articles: continue seen_articles.add(article_link) date_str = article.find_element(By.TAG_NAME, "time").get_attribute("datetime") publish_date = datetime.fromisoformat(date_str.replace("Z", "+00:00")) if publish_date < one_year_ago: # Stop scraping immediately driver.quit() print("Reached articles older than 1 year—stopping process.") exit() else: # Process the valid article print(f"Processing: {article.find_element(By.TAG_NAME, 'h2').text.strip()}") # Scroll to load more content driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # Adjust wait time based on page load speed
Key Notes for Both Tasks
- Timezone Consistency: Always normalize dates to the same timezone (e.g., UTC) to avoid incorrect comparisons. Reuters uses UTC in most
datetimeattributes, so align your cutoff date to UTC if needed. - Anti-Scraping Measures: Add delays between requests, use a valid
User-Agentheader, and consider rotating proxies if you hit rate limits. - Date Parsing Edge Cases: Some articles might use relative dates (e.g., "2 hours ago")—convert these to absolute dates using libraries like
dateutil.parser.
内容的提问来源于stack exchange,提问作者huy

