You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

项目征集信息自动Screen Scraping实现入门咨询

Web Scraping BMBF Funding Calls for Beginners: A Step-by-Step Guide

Hey Jonas, welcome to the world of web scraping! Since you're starting from scratch, let's break down exactly what you need to know to build your tool for extracting BMBF funding call info and exporting it to Excel.

1. Your Core Learning Path (入门方向)

First, focus on these foundational steps in order—they’ll build up to your end goal:

  • Understand HTML & Dev Tools: Spend 1-2 days learning basic HTML structure (tags, classes, IDs). Use your browser’s DevTools (F12 key) to inspect the elements on the BMBF pages you want to scrape. This lets you see exactly where the "Bekanntmachung" titles, deadlines, etc., are located in the page code.
  • Learn HTTP Requests: Understand how browsers fetch web pages. You’ll need to replicate this in code to get the page content programmatically.
  • Data Extraction: Master how to parse HTML and pull out specific pieces of data (like the funding call name or deadline) using a parsing library.
  • Data Handling & Export: Learn how to organize scraped data into a structured format (like a table) and export it to Excel.
  • Basic Anti-Scraping Etiquette: Always check the site’s robots.txt file to see if scraping is allowed. Keep your request rate slow (add small delays between requests) to avoid overwhelming the server—government sites often have strict rules here.

2. Best Programming Language for You: Python

Python is hands down the best choice for a beginner in this scenario. Here’s why, plus the key libraries you’ll use:

  • requests: This library lets you send HTTP requests to fetch the BMBF pages (replacing what your browser does when you visit a URL). It’s super easy to use for beginners.
  • BeautifulSoup4: A lightweight, beginner-friendly HTML parser. It lets you search for elements using their class names, IDs, or tag types—perfect since the BMBF funding calls have consistent formatting.
  • pandas: The go-to library for data handling in Python. Once you’ve scraped all your data, pandas lets you organize it into a table (DataFrame) and export it directly to an Excel file with just one line of code.

Quick Example Snippet (to give you a taste)

import requests
from bs4 import BeautifulSoup
import pandas as pd
from datetime import datetime, timedelta

# Define time range (last 6 months)
six_months_ago = datetime.now() - timedelta(days=180)

# List to store scraped data
funding_calls = []

# First, get the list page (you’ll need to find all Bekanntmachung links here)
list_url = "https://www.bmbf.de/foerderungen/"
response = requests.get(list_url)
soup = BeautifulSoup(response.text, "html.parser")

# Find all links to Bekanntmachung pages (adjust selector based on actual HTML)
call_links = soup.find_all("a", href=True)
for link in call_links:
    if "bekanntmachung" in link["href"]:
        call_url = f"https://www.bmbf.de{link['href']}"
        call_response = requests.get(call_url)
        call_soup = BeautifulSoup(call_response.text, "html.parser")
        
        # Extract data (adjust selectors to match actual BMBF page structure)
        title = call_soup.find("h1").text.strip()
        applicant = call_soup.find("div", class_="foerderempfaenger").text.strip()  # Example class
        funding_amount = call_soup.find("div", class_="foerderbetrag").text.strip()
        deadline = call_soup.find("div", class_="frist").text.strip()
        contact = call_soup.find("div", class_="kontakt").text.strip()
        
        # Parse deadline to check if it’s within last 6 months (adjust date format as needed)
        try:
            deadline_date = datetime.strptime(deadline, "%d.%m.%Y")
            if deadline_date >= six_months_ago:
                funding_calls.append({
                    "Title": title,
                    "Applicant": applicant,
                    "Funding Amount": funding_amount,
                    "Deadline": deadline,
                    "Contact": contact
                })
        except ValueError:
            # Skip if date format is unexpected
            continue

# Export to Excel
df = pd.DataFrame(funding_calls)
df.to_excel("bmbf_recent_funding_calls.xlsx", index=False)
print(f"Exported {len(funding_calls)} funding calls to Excel!")

Note: You’ll need to replace the class names (like foerderempfaenger) with the actual class/ID values you find using DevTools on the BMBF example page.

3. Useful Resources to Learn From

  • Official Library Docs:
    • requests: Read the quickstart guide to understand how to send requests and handle responses.
    • BeautifulSoup4: Go through the documentation’s "Quick Start" section to learn how to search for elements.
    • pandas: Check out the "Getting Started" guide, especially the part about exporting to Excel.
  • Beginner-Friendly Courses: Look for free Python scraping courses on platforms like Coursera or Udemy—search for "Python web scraping for beginners" to find structured learning paths.
  • Stack Overflow & GitHub:
    • Search Stack Overflow for questions like "How to scrape HTML tables with BeautifulSoup" or "Export pandas DataFrame to Excel" for solutions to common issues.
    • Browse GitHub for public projects tagged "web-scraping" or "government-data"—you’ll find examples of others scraping similar government sites that you can adapt.
  • HTML/CSS Basics: Spend a few hours on free learning sites to grasp enough HTML/CSS to understand how elements are structured on web pages.

内容的提问来源于stack exchange,提问作者Jonas Jung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 22:53:12