You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从AccuWeather网站抓取月度天气表格数据的技术请求

抓取AccuWeather班加罗尔月度天气表格数据的方案

Got it, let's walk through how to scrape the full monthly weather data for Bengaluru from that AccuWeather table. Here's a practical approach using Python with requests and BeautifulSoup—tools that are perfect for this kind of task:

1. 先准备依赖

First, install the required libraries if you haven't already:

pip install requests beautifulsoup4 pandas

2. 完整抓取代码

This script will fetch the page, parse the table, extract all daily weather details, and save them to a CSV file for easy use:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 模拟浏览器请求头,避免被网站反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# 目标页面链接
target_url = "https://www.accuweather.com/en/in/bengaluru/204108/month/204108?view=table"

try:
    # 发送请求获取页面内容
    response = requests.get(target_url, headers=headers)
    response.raise_for_status()  # 请求失败时抛出异常

    # 解析HTML内容
    soup = BeautifulSoup(response.text, "html.parser")

    # 定位到天气表格的主体部分
    weather_table_body = soup.find("tbody")
    if not weather_table_body:
        print("Oops, couldn't find the weather table—maybe the page structure changed?")
        exit()

    # 存储所有天气数据的列表
    all_weather_data = []

    # 遍历表格的每一行
    for day_row in weather_table_body.find_all("tr"):
        # 提取日期信息(星期+具体日期)
        date_header = day_row.find("th", scope="row")
        if not date_header:
            continue
        day_of_week = date_header.text.split()[0]
        date_str = date_header.find("time").text

        # 提取气温、降水量等列数据
        data_cells = day_row.find_all("td")
        if len(data_cells) < 3:
            continue  # 跳过数据不完整的行

        # 处理温度符号(替换HTML实体为℃)
        temp = data_cells[0].text.replace("&#176;", "°")
        # 提取当日降水量和月度累计降水量
        daily_precip = data_cells[1].text.strip()
        monthly_precip = data_cells[2].text.strip()

        # 将单天数据加入列表
        all_weather_data.append({
            "Day of Week": day_of_week,
            "Date": date_str,
            "Temperature (High/Low)": temp,
            "Daily Precipitation": daily_precip,
            "Monthly Total Precipitation": monthly_precip
        })

    # 转换为DataFrame并保存为CSV
    weather_df = pd.DataFrame(all_weather_data)
    weather_df.to_csv("bengaluru_monthly_weather.csv", index=False, encoding="utf-8")

    print("Done! All weather data saved to bengaluru_monthly_weather.csv")

except requests.exceptions.RequestException as e:
    print(f"Request failed: {e}")

3. 关键注意事项

  • 反爬应对:AccuWeather might block requests without a proper User-Agent, which is why we added the headers parameter. If you hit 403 errors, try adding a small delay between requests with time.sleep(1) or switch to a rotating proxy.
  • 页面结构变化:If the website updates its HTML later, you'll need to adjust the find()/find_all() selectors. Use your browser's dev tools to inspect the latest table structure.
  • 数据完整性:Rows with the pre class (like the one in your sample code) are future forecasts—this script handles them the same way as historical data, so you'll get the full month's data.

内容的提问来源于stack exchange,提问作者Ragu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:38:51