Python中如何基于字符串首元素对数据分组求平均值?
How to Average Data by Date in Python
Hey there! No need to apologize for your English at all—your problem is totally clear. Let's walk through how to group your data by date and calculate the averages for each metric.
First, let's adjust your existing code to collect all the data first (right now you're creating line_data but not saving it anywhere), then we'll group by date and compute averages.
Step-by-Step Solution Code
from collections import defaultdict # 1. Read and collect all valid data entries all_data = [] with open("Input2010_5a.txt", "r") as file: for line in file: # Note: Renamed `long` to `lon` to avoid conflict with Python's built-in `long` function raw_fields = line.strip().split("\t") if len(raw_fields) != 6: print(f"Skipping malformed line: {line.strip()}") continue date_str, lon, lat, depth, temp, sal = raw_fields try: # Convert all fields to floats date = float(date_str) line_data = [ date, float(lon), float(lat), float(depth), float(temp), float(sal) ] all_data.append(line_data) except ValueError as e: print(f"Skipping line with invalid number: {line.strip()}, error: {e}") # 2. Group data by date # We'll use a defaultdict to automatically create lists for new dates date_groups = defaultdict(list) for entry in all_data: date = entry[0] # Extract the metrics (everything except the date) metrics = entry[1:] date_groups[date].append(metrics) # 3. Calculate averages for each date averaged_results = [] for date, metrics_list in date_groups.items(): # Zip(*metrics_list) groups all values for the same metric together # e.g., all longitude values, all latitude values, etc. metric_averages = [sum(values) / len(values) for values in zip(*metrics_list)] # Combine date with its averages averaged_entry = [date] + metric_averages averaged_results.append(averaged_entry) # Optional: Print out the results to verify for res in averaged_results: print( f"Date: {res[0]} | " f"Avg Longitude: {res[1]:.2f} | " f"Avg Latitude: {res[2]:.2f} | " f"Avg Depth: {res[3]:.2f} | " f"Avg Temperature: {res[4]:.2f} | " f"Avg Salinity: {res[5]:.2f}" )
Key Notes:
- Avoiding built-in name conflicts: I renamed
longtolonbecauselongis a reserved function in Python—using it as a variable name can cause unexpected bugs later. - Error handling: Added checks for malformed lines (wrong number of fields) and invalid numeric values, so your script won't crash if there's messy data in the file.
- Grouping with
defaultdict: This makes it super easy to group entries by date without manually checking if a date key already exists in the dictionary. - Calculating averages: The
zip(*metrics_list)trick is handy here—it transposes your list of metrics, so you can compute the average for each metric (longitude, latitude, etc.) in one go.
Once you run this, averaged_results will hold a list where each entry is [date, avg_lon, avg_lat, avg_depth, avg_temp, avg_sal]—exactly what you need!
内容的提问来源于stack exchange,提问作者Alina Lerner
相关产品推荐
相关产品推荐

