处理大型DataFrame时遇TypeError:'method'对象不可下标访问
Hey there! Let's get this sorted out quickly. The error you're seeing (TypeError: 'method' object is not subscriptable) is a simple syntax issue, and we'll also tweak the code to generate the final1.csv to final24.csv files you want.
What's Causing the Error?
between_time is a method of pandas DataFrames, not a property. That means you need to call it with parentheses () instead of square brackets []. Your line df.between_time[day] is trying to subscript a method like a list, which Python doesn't allow.
Full Corrected Code
Here's the revised code that fixes the error and generates the hourly CSV files you need:
import pandas as pd # Read the CSV in chunks - removed parse_dates since we'll convert epoch manually reader = pd.read_csv( "D:/...path.../test.csv", names=["id_easy","ordinal", "latitude", "longitude","epoch",'weekday'], chunksize=n_rows, error_bad_lines=False ) # We'll pair each time range with an index (1-24) for the final filenames time_ranges = ( ('0:00', '1:00'),('1:00', '2:00'),('2:00', '3:00'),('3:00', '4:00'), ('4:00', '5:00'),('5:00', '6:00'), ('6:00', '7:00'),('7:00', '8:00'), ('8:00', '9:00'),('9:00', '10:00'),('10:00', '11:00'),('11:00', '12:00'), ('12:00', '13:00'),('13:00', '14:00'),('14:00', '15:00'),('15:00', '16:00'), ('16:00', '17:00'),('17:00', '18:00'), ('18:00', '19:00'),('19:00', '20:00'), ('20:00', '21:00'),('21:00', '22:00'),('22:00', '23:00'),('23:00', '00:00') ) for df in reader: if df.empty: continue # Skip empty chunks # Convert epoch (seconds since epoch) to datetime df['epoch'] = pd.to_datetime(df['epoch'], unit='s') # Set datetime as index for between_time to work df.set_index('epoch', inplace=True) # Loop through each time range with an index (1-24) for idx, (start_time, end_time) in enumerate(time_ranges, 1): # Correctly call between_time with start and end times hourly_df = df.between_time(start_time, end_time) if not hourly_df.empty: # Generate filename like final1.csv, final2.csv... hourly_df.to_csv( f'final{idx}.csv', index=False, header=False, mode='a' )
Key Changes Explained
- Fixed
between_timeCall: Changeddf.between_time[day]todf.between_time(start_time, end_time)to properly invoke the method with the start and end time parameters. - Filename Generation: Used
enumerate(time_ranges, 1)to get a 1-based index for each time range, so we createfinal1.csvtofinal24.csvinstead of using the time strings as filenames. - Optimized Datetime Handling: Removed
parse_dates=['epoch']fromread_csvsince we're explicitly converting the epoch column to datetime withunit='s'(avoids conflicting parsing). - Cleaner Empty Check: Used
continueto skip empty chunks early, making the code more readable.
This code will process your large CSV in chunks, split each chunk into hourly segments, and append each segment to the corresponding finalX.csv file.
内容的提问来源于stack exchange,提问作者Mamed

