多数据框特定时段点筛选及动物相近时间点经纬度距离计算
Hey there! Let's break down your two animal tracking data tasks and solve them with Python (using pandas and geopandas, the standard tools for this kind of work). Here's a step-by-step solution:
First, we'll make sure our timestamp columns are in datetime format (critical for time-based filtering), then apply boolean indexing to grab the time window we care about.
Assuming you have separate DataFrames for each animal (e.g., df_a for Animal A, df_b for Animal B, df_c for Animal C):
import pandas as pd # Convert timestamp columns to datetime type (do this for all DataFrames) df_a['timestamp'] = pd.to_datetime(df_a['timestamp']) df_b['timestamp'] = pd.to_datetime(df_b['timestamp']) df_c['timestamp'] = pd.to_datetime(df_c['timestamp']) # Define your target time window start_time = pd.to_datetime('2017-09-29 11:00:00') end_time = pd.to_datetime('2017-09-29 12:00:00') # Filter each DataFrame to the desired time range df_a_filtered = df_a[(df_a['timestamp'] >= start_time) & (df_a['timestamp'] <= end_time)] df_b_filtered = df_b[(df_b['timestamp'] >= start_time) & (df_b['timestamp'] <= end_time)] df_c_filtered = df_c[(df_c['timestamp'] >= start_time) & (df_c['timestamp'] <= end_time)]
This will give you cleaned DataFrames containing only the points from your specified time period.
Since your data is collected every 3 minutes but has small timestamp discrepancies (like Animal A's 11:10:08 entry), we'll use pd.merge_asof to match each Animal A entry with the closest timestamped entries from Animals B and C. Then we'll calculate distances using either geopandas (for precise projected distances) or a manual haversine formula.
Step 2.1: Match closest timestamps
merge_asof requires the right DataFrame to be sorted by the join key (timestamp), so we'll sort first:
# Sort all filtered DataFrames by timestamp df_a_filtered = df_a_filtered.sort_values('timestamp') df_b_filtered = df_b_filtered.sort_values('timestamp') df_c_filtered = df_c_filtered.sort_values('timestamp') # Match Animal A with the nearest Animal B entry merged_ab = pd.merge_asof( df_a_filtered, df_b_filtered, on='timestamp', direction='nearest', # Grabs the closest timestamp, regardless of earlier/later suffixes=('_a', '_b') # Add suffixes to avoid column name conflicts ) # Match the combined AB data with the nearest Animal C entry merged_abc = pd.merge_asof( merged_ab, df_c_filtered, on='timestamp', direction='nearest', suffixes=('', '_c') ) # Rename C's coordinates for clarity merged_abc.rename(columns={'long': 'long_c', 'lat': 'lat_c'}, inplace=True)
Step 2.2: Calculate meter-level distances
Option 1: Use Geopandas (most accurate for local data)
Geopandas lets us convert coordinates to a local projected CRS (like UTM, which uses meters as units) for precise distance calculations:
import geopandas as gpd from shapely.geometry import Point # Create geometry columns using WGS84 (global lat/lon CRS: EPSG:4326) merged_abc['point_a'] = gpd.points_from_xy(merged_abc['long_a'], merged_abc['lat_a'], crs='EPSG:4326') merged_abc['point_b'] = gpd.points_from_xy(merged_abc['long_b'], merged_abc['lat_b'], crs='EPSG:4326') merged_abc['point_c'] = gpd.points_from_xy(merged_abc['long_c'], merged_abc['lat_c'], crs='EPSG:4326') # Convert to UTM (auto-detects the correct UTM zone for your data) merged_abc = merged_abc.to_crs(merged_abc.estimate_utm_crs()) # Calculate distances in meters merged_abc['distance_a_b'] = merged_abc['point_a'].distance(merged_abc['point_b']) merged_abc['distance_a_c'] = merged_abc['point_a'].distance(merged_abc['point_c']) merged_abc['distance_b_c'] = merged_abc['point_b'].distance(merged_abc['point_c'])
Option 2: Manual Haversine Formula (no extra dependencies)
If you don't want to install geopandas, use the haversine formula to calculate spherical distances (accurate enough for most use cases):
import math def haversine(lon1, lat1, lon2, lat2): # Convert degrees to radians lon1, lat1, lon2, lat2 = map(math.radians, [lon1, lat1, lon2, lat2]) # Haversine formula to calculate great-circle distance dlon = lon2 - lon1 dlat = lat2 - lat1 a = math.sin(dlat/2)**2 + math.cos(lat1) * math.cos(lat2) * math.sin(dlon/2)**2 c = 2 * math.asin(math.sqrt(a)) earth_radius_m = 6371000 # Earth's radius in meters return c * earth_radius_m # Apply the function to calculate all pairwise distances merged_abc['distance_a_b'] = merged_abc.apply( lambda row: haversine(row['long_a'], row['lat_a'], row['long_b'], row['lat_b']), axis=1 ) merged_abc['distance_a_c'] = merged_abc.apply( lambda row: haversine(row['long_a'], row['lat_a'], row['long_c'], row['lat_c']), axis=1 ) merged_abc['distance_b_c'] = merged_abc.apply( lambda row: haversine(row['long_b'], row['lat_b'], row['long_c'], row['lat_c']), axis=1 )
After running this, your merged_abc DataFrame will have all the original data plus the calculated distances between each pair of animals at the closest matching timestamps.
内容的提问来源于stack exchange,提问作者kakadu

