You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取技术咨询:从指定外汇博客页面提取文章信息并生成CSV文件

Solution for Scraping GBP/USD Weekly Outlook Articles & Exporting to CSV

Let's fix up your scraper to get exactly the 11 GBP/USD weekly outlook articles you need, extract the right data, and export it to a properly formatted CSV. Here's a complete, targeted solution:

Key Issues Addressed

  • Your original code grabbed too many irrelevant links by only checking for "gbp" in URLs. We'll target actual article containers instead.
  • We'll extract the required fields (Year, Date, Title, Link) from each valid article.
  • We'll write everything to a CSV with your specified headers.

Complete Working Code

import httplib2
import csv
from datetime import datetime
from bs4 import BeautifulSoup

def scrape_gbp_usd_weekly_outlook(page_url):
    # Initialize HTTP client and fetch the target page
    http = httplib2.Http()
    status, response = http.request(page_url)
    
    # Parse the page HTML
    soup = BeautifulSoup(response, 'html.parser')
    
    # Target all blog post containers on the page
    article_containers = soup.find_all('article', class_='post')
    
    # Store extracted article data here
    article_data = []
    
    for article in article_containers:
        # Extract article title and direct link from the H2 header
        title_element = article.find('h2', class_='post-title').find('a')
        article_title = title_element.get_text(strip=True)
        article_link = title_element['href']
        
        # Filter only weekly outlook articles (adjust keyword if needed)
        if 'Weekly Outlook' not in article_title:
            continue
        
        # Extract and parse publication date from the <time> tag
        time_element = article.find('time', class_='post-date')
        date_string = time_element['datetime']
        parsed_date = datetime.strptime(date_string, '%Y-%m-%d')
        
        # Split date into Year and formatted Date (e.g., "October 16")
        article_year = parsed_date.year
        formatted_article_date = parsed_date.strftime('%B %d')
        
        # Add to our data list
        article_data.append({
            'Year': article_year,
            'Date': formatted_article_date,
            'Title': article_title,
            'Link': article_link
        })
    
    return article_data

def export_to_csv(data, filename='gbp_usd_weekly_outlook.csv'):
    # Define the exact CSV headers you requested
    csv_headers = ['Year', 'Date', 'Title', 'Link']
    
    # Write data to CSV with UTF-8 encoding
    with open(filename, 'w', newline='', encoding='utf-8') as csv_file:
        writer = csv.DictWriter(csv_file, fieldnames=csv_headers)
        writer.writeheader()
        writer.writerows(data)

if __name__ == '__main__':
    target_page = 'https://www.forexcrunch.com/category/forex-weekly-outlook/gbp-usd-outlook/page/24/'
    
    # Run the scraper
    scraped_articles = scrape_gbp_usd_weekly_outlook(target_page)
    
    # Verify we got the expected number of articles
    print(f"Successfully scraped {len(scraped_articles)} weekly outlook articles")
    
    # Export to CSV
    export_to_csv(scraped_articles)
    print("Data exported to gbp_usd_weekly_outlook.csv")

What This Code Does

  • Targeted Article Selection: We use soup.find_all('article', class_='post') to grab only the actual blog post containers, avoiding random navigation links.
  • Precise Filtering: We check for "Weekly Outlook" in the article title to ensure we only keep the content you want.
  • Clean Date Handling: The site uses a standard datetime attribute on the <time> tag, so we parse that to easily split into Year and formatted Date columns.
  • Proper CSV Export: Uses Python's built-in csv module to write data with your specified headers, and handles UTF-8 encoding to avoid character issues.

Setup & Notes

  1. Install required packages first:
    pip install httplib2 beautifulsoup4
    
  2. If the site's HTML structure changes (e.g., class names), use your browser's dev tools to inspect elements and update the selectors.
  3. Always respect the site's scraping policies (check robots.txt if unsure) to avoid being blocked.

内容的提问来源于stack exchange,提问作者backlog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:07:41