基于BeautifulSoup4爬取timeanddate.com天气数据的技术求助
Hey there! Since you already know the basics of pulling text from tags with BeautifulSoup, let's dive into grabbing those exact weather metrics you need—Feels Like, Visibility, Dew Point, Humidity, Wind, and Forecast—from the extended weather page. I'll use Mumbai as our example, but this will work for any city on that site once you adjust the country/place variables.
Step 1: Setup & Basic Page Parsing
First, let's get the page content parsed into a BeautifulSoup object. We'll use requests to fetch the page, so make sure you have both libraries installed (pip install requests beautifulsoup4).
import requests from bs4 import BeautifulSoup # Define your target location country = "india" place = "mumbai" quote_page = f"https://www.timeanddate.com/weather/{country}/{place}/ext" # Add headers to mimic a browser (helps avoid being blocked) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } # Fetch and parse the page response = requests.get(quote_page, headers=headers) soup = BeautifulSoup(response.text, 'html.parser')
Step 2: Extract Individual Current Weather Metrics
Let's break down how to grab each specific piece of data:
Feels Like
The "Feels Like" value lives in a div with class wts-value, right next to its label. We can find the label first, then grab its sibling element:
# Get Feels Like feels_like_label = soup.find('div', class_='wts-label', string='Feels Like') feels_like = feels_like_label.previous_sibling.text.strip() print(f"Feels Like: {feels_like}")
Visibility, Dew Point, Humidity, Wind
These four metrics are grouped in a table with class table--left table--inner-borders-vert. We can either loop through all rows to filter the ones we want, or grab each individually:
Option 1: Loop through the table (cleaner for multiple metrics)
# Grab the metrics table weather_table = soup.find('table', class_='table--left table--inner-borders-vert') # Extract only the metrics we care about target_metrics = ['Visibility', 'Dew Point', 'Humidity', 'Wind'] for row in weather_table.find_all('tr'): label = row.find('td').text.strip() value = row.find('td', class_='r').text.strip() if label in target_metrics: print(f"{label}: {value}")
Option 2: Grab each metric individually (good if you need only one or two)
# Get Visibility visibility = weather_table.find('td', string='Visibility').find_next_sibling('td').text.strip() # Get Dew Point dew_point = weather_table.find('td', string='Dew Point').find_next_sibling('td').text.strip() # Get Humidity humidity = weather_table.find('td', string='Humidity').find_next_sibling('td').text.strip() # Get Wind wind = weather_table.find('td', string='Wind').find_next_sibling('td').text.strip() print(f"Visibility: {visibility}\nDew Point: {dew_point}\nHumidity: {humidity}\nWind: {wind}")
Step 3: Extract the Forecast
The extended forecast is in a section with class bx bc-wt. Each day's forecast is wrapped in a div with class row-b. We'll extract key details (date, weather condition, high/low temps) and store them in a list of dictionaries for easy use:
# Extract forecast data forecast_section = soup.find('div', class_='bx bc-wt') forecast_list = [] for day in forecast_section.find_all('div', class_='row-b'): # Skip header rows (like "Date", "Weather", etc.) if not day.find('div', class_='date'): continue date = day.find('div', class_='date').text.strip() weather_condition = day.find('div', class_='small').text.strip() high_temp = day.find('div', class_='hi').text.strip() low_temp = day.find('div', class_='lo').text.strip() # Add to our forecast list forecast_list.append({ 'date': date, 'condition': weather_condition, 'high_temp': high_temp, 'low_temp': low_temp }) # Print the forecast print("\nExtended Forecast:") for entry in forecast_list: print(f"{entry['date']}: {entry['condition']} | High: {entry['high_temp']}, Low: {entry['low_temp']}")
A Quick Heads-Up
- Page Structure Changes: Websites often tweak their HTML, so if this code stops working later, check the page source to confirm class names or element positions haven't changed.
- Error Handling: For robustness, wrap your find calls in
try-exceptblocks to handle cases where elements might be missing (e.g., if the site loads incorrectly).
Hope this helps you level up your BeautifulSoup skills—happy scraping!
内容的提问来源于stack exchange,提问作者GFTW

