You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于网页抓取的天气数据函数,如何遍历万行数据集批量生成DataFrame?

Solution: Generate Full Weather Dataset by Traversing Your 10k-Row Dataset

First, let's tweak your existing get_weather function to fix a few small gaps that'll make it work reliably for bulk processing:

  • It wasn't returning the DataFrame you created
  • If your latitude/longitude/unix parameters are numeric (not strings), concatenating them into the URL will throw an error
  • We'll add basic error handling to avoid crashes if an API request fails

Here's the updated function:

import pandas as pd
import requests
import time

def get_weather(latitude, longitude, unix):
    # Convert numeric parameters to strings for valid URL formatting
    lat_str = str(latitude)
    lon_str = str(longitude)
    unix_str = str(unix)
    
    url = f"https://dark-sky.p.rapidapi.com/{lat_str},{lon_str},{unix_str}"
    headers = {
        'x-rapidapi-key': "xxxxxxxxxxxxxxMYKEYxxxxxxxxxxxxxxx",
        'x-rapidapi-host': "dark-sky.p.rapidapi.com"
    }
    
    try:
        response = requests.request("GET", url, headers=headers)
        response.raise_for_status()  # Trigger error for HTTP 4xx/5xx responses
        data = response.json()
        weather = data['currently']
        return pd.DataFrame(weather, index=[0])
    except Exception as e:
        print(f"Failed to fetch weather for ({lat_str}, {lon_str}) at {unix_str}: {str(e)}")
        # Return empty DataFrame with expected columns to keep dataset consistency
        return pd.DataFrame(columns=['time', 'summary', 'icon', 'precipIntensity', 'precipProbability',
                                     'precipType', 'temperature', 'apparentTemperature', 'dewPoint',
                                     'humidity', 'pressure', 'windSpeed', 'windBearing', 'cloudCover',
                                     'uvIndex', 'visibility'])

Next, let's handle the bulk processing of your 10k-row dataset. We'll assume your source data is stored in a pandas DataFrame called source_df, with columns latitude, longitude, and unix_time (adjust these names to match your actual dataset).

Method 1: Iterative Approach (Simple, Easy to Debug)

This method loops through each row, fetches the weather data, and builds your full dataset step by step:

# Initialize an empty list to store individual weather DataFrames
weather_results = []

# Loop through every row in your source dataset
for _, row in source_df.iterrows():
    weather_df = get_weather(row['latitude'], row['longitude'], row['unix_time'])
    weather_results.append(weather_df)
    # Add a small delay to avoid hitting API rate limits (adjust based on your RapidAPI plan)
    time.sleep(0.1)

# Combine all individual DataFrames into one full weather dataset
full_weather_dataset = pd.concat(weather_results, ignore_index=True)

# Optional: Merge with your original source data if you want to keep all columns together
final_combined_dataset = pd.concat([source_df.reset_index(drop=True), full_weather_dataset], axis=1)

Method 2: Parallel Processing (Faster for Large Datasets)

If you want to speed up the process, you can use parallel threads (just make sure you don't exceed your API's rate limits):

from concurrent.futures import ThreadPoolExecutor

# Helper function to process a single row
def process_single_row(row_tuple):
    _, row = row_tuple
    weather_df = get_weather(row['latitude'], row['longitude'], row['unix_time'])
    time.sleep(0.1)  # Still respect rate limits
    return weather_df

# Use a thread pool to process rows in parallel
with ThreadPoolExecutor(max_workers=5) as executor:
    weather_results = list(executor.map(process_single_row, source_df.iterrows()))

# Combine results into the final dataset
full_weather_dataset = pd.concat(weather_results, ignore_index=True)
final_combined_dataset = pd.concat([source_df.reset_index(drop=True), full_weather_dataset], axis=1)

Key Things to Keep in Mind:

  • API Rate Limits: RapidAPI and Dark Sky enforce rate limits—check your plan's limits and adjust the time.sleep() duration to avoid getting blocked.
  • Error Handling: The updated get_weather function catches failures and returns an empty row with the correct columns, so a single bad request won't break your entire dataset.
  • Data Consistency: Using ignore_index=True when concatenating ensures your final dataset has a clean, continuous index.

内容的提问来源于stack exchange,提问作者Alex Zieky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:57:24