Python按月日分组求最值并生成独立DataFrame(规避NumPy)
Got it, let's fix this up step by step—no NumPy required, just pure pandas. The key issues you might have hit are probably untyped data (your milesrun is stored as strings!) or incorrect date parsing, so we'll address those first.
Step 1: Convert Raw Data to DataFrame & Clean Columns
First, let's get your dictionary into a pandas DataFrame and fix the data types to ensure proper calculations:
import pandas as pd # Your raw daily data data1 = {'Date': {1: '01-01-2001', 2: '01-01-2002', 3: '01-01-2003', 4: '01-01-2004', 5: '01-01-2005', 6: '01-01-2006', 7: '01-01-2007', 8: '01-01-2008', 9: '01-01-2009', 10: '01-01-2010' }, 'milesrun': {1: '15', 2: '21', 3: '19', 4: '22', 5: '16', 6: '13', 7: '22', 8: '24', 9: '17', 10: '18'}} # Convert dictionary to pandas DataFrame df = pd.DataFrame(data1) # Parse dates correctly (adjust format to '%m-%d-%Y' if your dates are mm-dd-yyyy instead of dd-mm-yyyy) df['Date'] = pd.to_datetime(df['Date'], format='%d-%m-%Y') # Extract month-day string as grouping key (use '%d-%m' if you prefer day-month order) df['mth-date'] = df['Date'].dt.strftime('%m-%d') # Convert milesrun from string to integer (critical for accurate min/max calculations!) df['milesrun'] = df['milesrun'].astype(int)
Step 2: Group & Split into Min/Max DataFrames
Now we can group by the mth-date column, compute both stats in one pass, then split them into separate DataFrames matching your required structure:
# Group by month-day and calculate min/max miles in a single aggregation grouped_stats = df.groupby('mth-date')['milesrun'].agg(['min', 'max']) # Create DataFrame for minimum values (with your required column names) min_miles_df = grouped_stats[['min']].reset_index().rename(columns={'min': 'value'}) # Create DataFrame for maximum values max_miles_df = grouped_stats[['max']].reset_index().rename(columns={'max': 'value'})
Step 3: Check the Output
If you print the results, you'll get exactly what you need:
min_miles_dfhas columnsmth-date(e.g., '01-01' for January 1st) andvalue(the lowest miles run on that month-day across all years)max_miles_dffollows the same structure but contains the highest miles run for each month-day
For your sample data, the min value is 13 and max is 24, both tied to '01-01' since all entries are for January 1st.
Why This Works (And What You Might Have Missed)
- String-to-Numeric Conversion: Your original
milesrunvalues are stored as strings—grouping without converting them would lead to lexicographical comparison (e.g., '15' > '2') which gives wrong min/max results. - Reliable Date Parsing: Using
pd.to_datetimeensures we correctly extract the month-day part, avoiding errors from manual string slicing that can break if date formats vary. - No NumPy Dependency: All operations use pandas' built-in grouping and aggregation functions, so you don't need to import any additional libraries.
内容的提问来源于stack exchange,提问作者rajeev

