You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用循环对DataFrame中distance与bearing列组合进行1000次重采样?

Hey there! Let's walk through how to solve this problem using Python—it's a flexible tool for this kind of data generation task, and I'll keep the example easy to adapt to your actual data once you have it ready.

Step 1: Set Up Dependencies & Data Structure

First, we'll use numpy for robust random number generation and pandas (optional but super handy for tabular data) to work with your dataset. If you don't have these installed, run this in your terminal:

pip install numpy pandas
Step 2: Mock Your Original Data (Swap with Yours Later!)

Since you haven't shared your actual data yet, let's create a sample 16-row dataset you can replace once you're ready:

import numpy as np
import pandas as pd

# Sample original data (16 rows of distance + bearing)
original_data = pd.DataFrame({
    'distance': np.random.uniform(10, 100, 16),  # Replace with your real distances
    'bearing': np.random.uniform(0, 360, 16)     # Replace with your real bearings
})
Step 3: Define a Function to Generate 16 New Random Pairs

The core task is creating new pairs based on your original values. Two common approaches fit this need:

  • Gaussian Disturbance: Add small, natural-sounding noise to each original value (great if you want values close to the original)
  • Range-Based Randomization: Pick a random value within a fixed range around each original value

Let's start with the Gaussian approach—you can swap it out if you need something else:

def generate_new_group(original_df, distance_std=5, bearing_std=10):
    """
    Generate 16 new (distance, bearing) pairs using original data as a baseline.
    
    Args:
        original_df: DataFrame with 'distance' and 'bearing' columns
        distance_std: How much variation to allow in distance (adjust as needed)
        bearing_std: How much variation to allow in bearing (adjust as needed)
    
    Returns:
        DataFrame with 16 new randomized pairs
    """
    # Add Gaussian noise to distances
    new_distances = np.random.normal(original_df['distance'], distance_std)
    # Add noise to bearings, then wrap to 0-360 degrees (since bearings are angles)
    new_bearings = np.mod(np.random.normal(original_df['bearing'], bearing_std), 360)
    
    return pd.DataFrame({
        'new_distance': new_distances,
        'new_bearing': new_bearings
    })

If you prefer range-based randomization (e.g., ±10% of original distance, ±15 degrees of bearing), use this version instead:

def generate_new_group_range(original_df, distance_range_pct=0.1, bearing_range=15):
    new_distances = np.random.uniform(
        original_df['distance'] * (1 - distance_range_pct),
        original_df['distance'] * (1 + distance_range_pct)
    )
    new_bearings = np.mod(
        np.random.uniform(
            original_df['bearing'] - bearing_range,
            original_df['bearing'] + bearing_range
        ),
        360
    )
    return pd.DataFrame({
        'new_distance': new_distances,
        'new_bearing': new_bearings
    })
Step 4: Run the Loop 1000 Times

Now we'll generate 1000 groups (each with 16 pairs) and store them all. We'll add a run_id to track which group each pair belongs to:

num_runs = 1000
all_results = []

for run_id in range(num_runs):
    # Use either generate_new_group or generate_new_group_range here
    new_group = generate_new_group(original_data)
    new_group['run_id'] = run_id  # Label which iteration this group comes from
    all_results.append(new_group)

# Combine all results into one easy-to-manage DataFrame (optional)
combined_results = pd.concat(all_results, ignore_index=True)
Step 5: Save or Use Your Results

If you want to save the full dataset for later analysis, export it to a CSV:

combined_results.to_csv('randomized_bearing_distance.csv', index=False)
Quick Adaptation Tips
  • Load your real data: Once you have your dataset, replace the mock original_data with pd.read_csv('your_data_file.csv').
  • Tweak randomness: Adjust the distance_std/bearing_std or range parameters to control how much variation you want in your new pairs.
  • R alternative: If you prefer R, the logic is identical—use dplyr for data handling and rnorm()/runif() for random generation.

内容的提问来源于stack exchange,提问作者MMaor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:57:52