基于网页抓取的天气数据函数,如何遍历万行数据集批量生成DataFrame?
Solution: Generate Full Weather Dataset by Traversing Your 10k-Row Dataset
First, let's tweak your existing get_weather function to fix a few small gaps that'll make it work reliably for bulk processing:
- It wasn't returning the DataFrame you created
- If your latitude/longitude/unix parameters are numeric (not strings), concatenating them into the URL will throw an error
- We'll add basic error handling to avoid crashes if an API request fails
Here's the updated function:
import pandas as pd import requests import time def get_weather(latitude, longitude, unix): # Convert numeric parameters to strings for valid URL formatting lat_str = str(latitude) lon_str = str(longitude) unix_str = str(unix) url = f"https://dark-sky.p.rapidapi.com/{lat_str},{lon_str},{unix_str}" headers = { 'x-rapidapi-key': "xxxxxxxxxxxxxxMYKEYxxxxxxxxxxxxxxx", 'x-rapidapi-host': "dark-sky.p.rapidapi.com" } try: response = requests.request("GET", url, headers=headers) response.raise_for_status() # Trigger error for HTTP 4xx/5xx responses data = response.json() weather = data['currently'] return pd.DataFrame(weather, index=[0]) except Exception as e: print(f"Failed to fetch weather for ({lat_str}, {lon_str}) at {unix_str}: {str(e)}") # Return empty DataFrame with expected columns to keep dataset consistency return pd.DataFrame(columns=['time', 'summary', 'icon', 'precipIntensity', 'precipProbability', 'precipType', 'temperature', 'apparentTemperature', 'dewPoint', 'humidity', 'pressure', 'windSpeed', 'windBearing', 'cloudCover', 'uvIndex', 'visibility'])
Next, let's handle the bulk processing of your 10k-row dataset. We'll assume your source data is stored in a pandas DataFrame called source_df, with columns latitude, longitude, and unix_time (adjust these names to match your actual dataset).
Method 1: Iterative Approach (Simple, Easy to Debug)
This method loops through each row, fetches the weather data, and builds your full dataset step by step:
# Initialize an empty list to store individual weather DataFrames weather_results = [] # Loop through every row in your source dataset for _, row in source_df.iterrows(): weather_df = get_weather(row['latitude'], row['longitude'], row['unix_time']) weather_results.append(weather_df) # Add a small delay to avoid hitting API rate limits (adjust based on your RapidAPI plan) time.sleep(0.1) # Combine all individual DataFrames into one full weather dataset full_weather_dataset = pd.concat(weather_results, ignore_index=True) # Optional: Merge with your original source data if you want to keep all columns together final_combined_dataset = pd.concat([source_df.reset_index(drop=True), full_weather_dataset], axis=1)
Method 2: Parallel Processing (Faster for Large Datasets)
If you want to speed up the process, you can use parallel threads (just make sure you don't exceed your API's rate limits):
from concurrent.futures import ThreadPoolExecutor # Helper function to process a single row def process_single_row(row_tuple): _, row = row_tuple weather_df = get_weather(row['latitude'], row['longitude'], row['unix_time']) time.sleep(0.1) # Still respect rate limits return weather_df # Use a thread pool to process rows in parallel with ThreadPoolExecutor(max_workers=5) as executor: weather_results = list(executor.map(process_single_row, source_df.iterrows())) # Combine results into the final dataset full_weather_dataset = pd.concat(weather_results, ignore_index=True) final_combined_dataset = pd.concat([source_df.reset_index(drop=True), full_weather_dataset], axis=1)
Key Things to Keep in Mind:
- API Rate Limits: RapidAPI and Dark Sky enforce rate limits—check your plan's limits and adjust the
time.sleep()duration to avoid getting blocked. - Error Handling: The updated
get_weatherfunction catches failures and returns an empty row with the correct columns, so a single bad request won't break your entire dataset. - Data Consistency: Using
ignore_index=Truewhen concatenating ensures your final dataset has a clean, continuous index.
内容的提问来源于stack exchange,提问作者Alex Zieky
相关产品推荐
相关产品推荐

