You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python按月日分组求最值并生成独立DataFrame(规避NumPy)

Solution: Group by Month-Day & Get Min/Max Miles (Pandas-Only, No NumPy)

Got it, let's fix this up step by step—no NumPy required, just pure pandas. The key issues you might have hit are probably untyped data (your milesrun is stored as strings!) or incorrect date parsing, so we'll address those first.

Step 1: Convert Raw Data to DataFrame & Clean Columns

First, let's get your dictionary into a pandas DataFrame and fix the data types to ensure proper calculations:

import pandas as pd

# Your raw daily data
data1 = {'Date': {1: '01-01-2001', 2: '01-01-2002', 3: '01-01-2003', 4: '01-01-2004', 5: '01-01-2005', 6: '01-01-2006', 7: '01-01-2007', 8: '01-01-2008', 9: '01-01-2009', 10: '01-01-2010' }, 'milesrun': {1: '15', 2: '21', 3: '19', 4: '22', 5: '16', 6: '13', 7: '22', 8: '24', 9: '17', 10: '18'}}

# Convert dictionary to pandas DataFrame
df = pd.DataFrame(data1)

# Parse dates correctly (adjust format to '%m-%d-%Y' if your dates are mm-dd-yyyy instead of dd-mm-yyyy)
df['Date'] = pd.to_datetime(df['Date'], format='%d-%m-%Y')

# Extract month-day string as grouping key (use '%d-%m' if you prefer day-month order)
df['mth-date'] = df['Date'].dt.strftime('%m-%d')

# Convert milesrun from string to integer (critical for accurate min/max calculations!)
df['milesrun'] = df['milesrun'].astype(int)

Step 2: Group & Split into Min/Max DataFrames

Now we can group by the mth-date column, compute both stats in one pass, then split them into separate DataFrames matching your required structure:

# Group by month-day and calculate min/max miles in a single aggregation
grouped_stats = df.groupby('mth-date')['milesrun'].agg(['min', 'max'])

# Create DataFrame for minimum values (with your required column names)
min_miles_df = grouped_stats[['min']].reset_index().rename(columns={'min': 'value'})

# Create DataFrame for maximum values
max_miles_df = grouped_stats[['max']].reset_index().rename(columns={'max': 'value'})

Step 3: Check the Output

If you print the results, you'll get exactly what you need:

  • min_miles_df has columns mth-date (e.g., '01-01' for January 1st) and value (the lowest miles run on that month-day across all years)
  • max_miles_df follows the same structure but contains the highest miles run for each month-day

For your sample data, the min value is 13 and max is 24, both tied to '01-01' since all entries are for January 1st.

Why This Works (And What You Might Have Missed)

  • String-to-Numeric Conversion: Your original milesrun values are stored as strings—grouping without converting them would lead to lexicographical comparison (e.g., '15' > '2') which gives wrong min/max results.
  • Reliable Date Parsing: Using pd.to_datetime ensures we correctly extract the month-day part, avoiding errors from manual string slicing that can break if date formats vary.
  • No NumPy Dependency: All operations use pandas' built-in grouping and aggregation functions, so you don't need to import any additional libraries.

内容的提问来源于stack exchange,提问作者rajeev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:45:01