网页抓取技术咨询:从指定外汇博客页面提取文章信息并生成CSV文件
Solution for Scraping GBP/USD Weekly Outlook Articles & Exporting to CSV
Let's fix up your scraper to get exactly the 11 GBP/USD weekly outlook articles you need, extract the right data, and export it to a properly formatted CSV. Here's a complete, targeted solution:
Key Issues Addressed
- Your original code grabbed too many irrelevant links by only checking for "gbp" in URLs. We'll target actual article containers instead.
- We'll extract the required fields (Year, Date, Title, Link) from each valid article.
- We'll write everything to a CSV with your specified headers.
Complete Working Code
import httplib2 import csv from datetime import datetime from bs4 import BeautifulSoup def scrape_gbp_usd_weekly_outlook(page_url): # Initialize HTTP client and fetch the target page http = httplib2.Http() status, response = http.request(page_url) # Parse the page HTML soup = BeautifulSoup(response, 'html.parser') # Target all blog post containers on the page article_containers = soup.find_all('article', class_='post') # Store extracted article data here article_data = [] for article in article_containers: # Extract article title and direct link from the H2 header title_element = article.find('h2', class_='post-title').find('a') article_title = title_element.get_text(strip=True) article_link = title_element['href'] # Filter only weekly outlook articles (adjust keyword if needed) if 'Weekly Outlook' not in article_title: continue # Extract and parse publication date from the <time> tag time_element = article.find('time', class_='post-date') date_string = time_element['datetime'] parsed_date = datetime.strptime(date_string, '%Y-%m-%d') # Split date into Year and formatted Date (e.g., "October 16") article_year = parsed_date.year formatted_article_date = parsed_date.strftime('%B %d') # Add to our data list article_data.append({ 'Year': article_year, 'Date': formatted_article_date, 'Title': article_title, 'Link': article_link }) return article_data def export_to_csv(data, filename='gbp_usd_weekly_outlook.csv'): # Define the exact CSV headers you requested csv_headers = ['Year', 'Date', 'Title', 'Link'] # Write data to CSV with UTF-8 encoding with open(filename, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=csv_headers) writer.writeheader() writer.writerows(data) if __name__ == '__main__': target_page = 'https://www.forexcrunch.com/category/forex-weekly-outlook/gbp-usd-outlook/page/24/' # Run the scraper scraped_articles = scrape_gbp_usd_weekly_outlook(target_page) # Verify we got the expected number of articles print(f"Successfully scraped {len(scraped_articles)} weekly outlook articles") # Export to CSV export_to_csv(scraped_articles) print("Data exported to gbp_usd_weekly_outlook.csv")
What This Code Does
- Targeted Article Selection: We use
soup.find_all('article', class_='post')to grab only the actual blog post containers, avoiding random navigation links. - Precise Filtering: We check for "Weekly Outlook" in the article title to ensure we only keep the content you want.
- Clean Date Handling: The site uses a standard
datetimeattribute on the<time>tag, so we parse that to easily split into Year and formatted Date columns. - Proper CSV Export: Uses Python's built-in
csvmodule to write data with your specified headers, and handles UTF-8 encoding to avoid character issues.
Setup & Notes
- Install required packages first:
pip install httplib2 beautifulsoup4 - If the site's HTML structure changes (e.g., class names), use your browser's dev tools to inspect elements and update the selectors.
- Always respect the site's scraping policies (check
robots.txtif unsure) to avoid being blocked.
内容的提问来源于stack exchange,提问作者backlog
相关产品推荐
相关产品推荐

