You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Django中从第三方网站获取网站数据并提取指定整数值?(以谷歌温度页面为例)

Web Scraping Integer Values (Like Temperature) in Django for Beginners

Hey there! Let's walk through how to scrape data like temperature from third-party sites (like Google's search results) using Django, since you're starting from scratch. We'll use two core tools for this: requests (to fetch web pages) and BeautifulSoup (to parse HTML content).

Step 1: Install Required Packages

First, install the tools we need. Open your terminal and run:

pip install requests beautifulsoup4

Step 2: Django View Code (views.py)

Here's a complete example of a view that fetches today's temperature from Google's search results. I'll add comments to explain each part:

from django.shortcuts import render
from django.http import HttpResponse
import requests
from bs4 import BeautifulSoup

def get_temperature(request):
    # URL for Google search query "today temperature"
    search_url = "https://www.google.com/search?q=today+temperature"
    
    # Headers to mimic a real browser (avoids being blocked by Google)
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    try:
        # Send GET request to fetch the page
        response = requests.get(search_url, headers=headers)
        response.raise_for_status()  # Raise error if request fails (e.g., 404, 500)
        
        # Parse the HTML content with BeautifulSoup
        soup = BeautifulSoup(response.text, 'html.parser')
        
        # Find the temperature element (Google's class may change, check with dev tools!)
        # As of 2024, this class targets the temperature value
        temp_element = soup.find(class_="BNeawe iBp4i AP7Wnd")
        
        if temp_element:
            # Extract the text and clean it (e.g., "25°C" becomes "25")
            temp_text = temp_element.text.strip()
            # Extract the integer part
            temperature = int(''.join(filter(str.isdigit, temp_text)))
            
            # Return the result as an HttpResponse or render a template
            return HttpResponse(f"Today's temperature is {temperature}°C")
        else:
            return HttpResponse("Could not find temperature data on the page.")
    
    except requests.exceptions.RequestException as e:
        # Handle request errors (e.g., no internet, blocked by Google)
        return HttpResponse(f"Error fetching data: {str(e)}")

Key Notes for the Code:

  • Headers: Google blocks requests without a proper User-Agent (it detects non-browser requests). The header above mimics a Chrome browser.
  • Element Selection: To find the right class/id for the temperature, right-click the temperature on Google's page → "Inspect" to see the HTML structure. If Google updates their UI, you'll need to update this selector.
  • Error Handling: We use try/except to catch common issues like network errors or missing elements, so your Django app doesn't crash unexpectedly.

Step 3: Map the View to a URL (urls.py)

Don't forget to add a URL pattern so you can access this view:

from django.urls import path
from . import views

urlpatterns = [
    path('temperature/', views.get_temperature, name='get_temperature'),
]

Learning Resources for Beginners

Since you're new to web scraping, here are some great places to learn (you can search these directly):

  • Requests Official Documentation: Covers everything about sending HTTP requests, handling responses, and customizing headers.
  • BeautifulSoup Official Documentation: Step-by-step guides to parsing HTML, finding elements by class, id, tag, and more.
  • Django Official Views Documentation: Reinforce your understanding of how Django views work, since scraping will live inside views in your app.
  • Stack Overflow's "web-scraping" Tag: Browse questions and answers about common scraping issues—you'll find tons of real-world examples.
  • Real Python's Web Scraping 101 Series: A beginner-friendly series that walks through scraping basics with Python, including how to handle dynamic content (if you need to scrape sites with JavaScript later).

Important Reminders

  • Respect robots.txt: Most sites have a robots.txt file (e.g., https://www.google.com/robots.txt) that tells you which parts of the site you can scrape. Always check this first.
  • Avoid Overloading Servers: Don't send too many requests in a short time—this can get your IP blocked. Add delays if you're scraping multiple pages.
  • Site Changes: Third-party sites often update their HTML structure, so your selector may break over time. You'll need to revisit and adjust the code if that happens.

内容的提问来源于stack exchange,提问作者Kanchon Gharami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 03:22:46