如何在Django中从第三方网站获取网站数据并提取指定整数值?(以谷歌温度页面为例)
Web Scraping Integer Values (Like Temperature) in Django for Beginners
Hey there! Let's walk through how to scrape data like temperature from third-party sites (like Google's search results) using Django, since you're starting from scratch. We'll use two core tools for this: requests (to fetch web pages) and BeautifulSoup (to parse HTML content).
Step 1: Install Required Packages
First, install the tools we need. Open your terminal and run:
pip install requests beautifulsoup4
Step 2: Django View Code (views.py)
Here's a complete example of a view that fetches today's temperature from Google's search results. I'll add comments to explain each part:
from django.shortcuts import render from django.http import HttpResponse import requests from bs4 import BeautifulSoup def get_temperature(request): # URL for Google search query "today temperature" search_url = "https://www.google.com/search?q=today+temperature" # Headers to mimic a real browser (avoids being blocked by Google) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } try: # Send GET request to fetch the page response = requests.get(search_url, headers=headers) response.raise_for_status() # Raise error if request fails (e.g., 404, 500) # Parse the HTML content with BeautifulSoup soup = BeautifulSoup(response.text, 'html.parser') # Find the temperature element (Google's class may change, check with dev tools!) # As of 2024, this class targets the temperature value temp_element = soup.find(class_="BNeawe iBp4i AP7Wnd") if temp_element: # Extract the text and clean it (e.g., "25°C" becomes "25") temp_text = temp_element.text.strip() # Extract the integer part temperature = int(''.join(filter(str.isdigit, temp_text))) # Return the result as an HttpResponse or render a template return HttpResponse(f"Today's temperature is {temperature}°C") else: return HttpResponse("Could not find temperature data on the page.") except requests.exceptions.RequestException as e: # Handle request errors (e.g., no internet, blocked by Google) return HttpResponse(f"Error fetching data: {str(e)}")
Key Notes for the Code:
- Headers: Google blocks requests without a proper
User-Agent(it detects non-browser requests). The header above mimics a Chrome browser. - Element Selection: To find the right class/id for the temperature, right-click the temperature on Google's page → "Inspect" to see the HTML structure. If Google updates their UI, you'll need to update this selector.
- Error Handling: We use
try/exceptto catch common issues like network errors or missing elements, so your Django app doesn't crash unexpectedly.
Step 3: Map the View to a URL (urls.py)
Don't forget to add a URL pattern so you can access this view:
from django.urls import path from . import views urlpatterns = [ path('temperature/', views.get_temperature, name='get_temperature'), ]
Learning Resources for Beginners
Since you're new to web scraping, here are some great places to learn (you can search these directly):
- Requests Official Documentation: Covers everything about sending HTTP requests, handling responses, and customizing headers.
- BeautifulSoup Official Documentation: Step-by-step guides to parsing HTML, finding elements by class, id, tag, and more.
- Django Official Views Documentation: Reinforce your understanding of how Django views work, since scraping will live inside views in your app.
- Stack Overflow's "web-scraping" Tag: Browse questions and answers about common scraping issues—you'll find tons of real-world examples.
- Real Python's Web Scraping 101 Series: A beginner-friendly series that walks through scraping basics with Python, including how to handle dynamic content (if you need to scrape sites with JavaScript later).
Important Reminders
- Respect
robots.txt: Most sites have arobots.txtfile (e.g.,https://www.google.com/robots.txt) that tells you which parts of the site you can scrape. Always check this first. - Avoid Overloading Servers: Don't send too many requests in a short time—this can get your IP blocked. Add delays if you're scraping multiple pages.
- Site Changes: Third-party sites often update their HTML structure, so your selector may break over time. You'll need to revisit and adjust the code if that happens.
内容的提问来源于stack exchange,提问作者Kanchon Gharami
相关产品推荐
相关产品推荐

