无法对Pandas DataFrame排序及read_csv参数相关疑问求助
Hey there, let's break down your questions one by one and fix that frustrating KeyError: 'timestamp' issue!
1. Does your read_csv call meet your requirements?
For the most part, yes—but we can make it more robust. Let's walk through your parameters:
usecols=[0,1,2,3,4,5,6]: Correctly selects the first 7 columns of your CSV (timestamp to Volume), but using column names instead of indices is safer (more on that next).header=0: Properly tells pandas that the first row of your CSV is the header (column names).index_col=0: Sets the first column (timestamp) as the DataFrame's index—great for time-series data, as it aligns with the old example's approach.parse_dates=True: Automatically parses your index (timestamp) into datetime objects, which is exactly what you want for sorting and time-based operations.
2. Can you select columns by name instead of index?
Absolutely! This is actually better practice because it avoids errors if your CSV's column order ever changes. Just replace the numeric usecols list with a list of column names matching your CSV:
usecols=['timestamp', 'open', 'high', 'low', 'close', 'adjusted_close', 'Volume']
3. What does the names parameter correspond to?
This is a common point of confusion:
- If
header=0(your case, since your CSV has a header row),namesreplaces the existing column names from the CSV. For example, if your CSV's 7th column is namedVolume, settingnames=[..., 'volume']will rename that column to lowercasevolume. - If your CSV had no header row, you'd use
header=Noneandnamesto define column names from scratch. - In your original code, since you set
index_col=0, thetimestampcolumn becomes the index—not a regular column—so renaming it withnamesdoesn't affect the index label (though it's unnecessary here).
4. Is sort_values(by='timestamp') equivalent to the old .sort()?
Almost, but with a key difference:
- In older pandas (Python 2.x era),
.sort()sorted the DataFrame by its index by default. Since the old example used the datetime column as the index,.sort()was sorting by date. - In your pandas 0.22.0,
.sort()was replaced by two methods:.sort_index(): Sorts by the DataFrame's index (this is the direct equivalent of the old.sort()in your case, since timestamp is your index)..sort_values(): Sorts by one or more column values. Since you settimestampas the index, it's no longer a regular column—hence the KeyError when you try to reference it here.
5. How to fix the KeyError: 'timestamp'?
The root cause is that you're trying to sort by timestamp as a column, but you set it as the DataFrame's index with index_col=0. Here are two easy fixes:
Fix 1: Use .sort_index() (recommended for your setup)
Since timestamp is your index, sorting by the index is the most efficient approach and matches the old example's behavior:
self.symbol_data[s] = pd.read_csv( os.path.join(self.csv_dir, '%s.csv' % s), usecols=['timestamp', 'open', 'high', 'low', 'close', 'adjusted_close', 'Volume'], header=0, index_col='timestamp', # Use column name instead of index for clarity parse_dates=True ).sort_index()
Fix 2: Keep timestamp as a regular column
If you prefer to keep timestamp as a column instead of the index, remove index_col=0 and explicitly parse it as a datetime:
self.symbol_data[s] = pd.read_csv( os.path.join(self.csv_dir, '%s.csv' % s), usecols=['timestamp', 'open', 'high', 'low', 'close', 'adjusted_close', 'Volume'], header=0, parse_dates=['timestamp'] ).sort_values(by='timestamp')
Bonus Troubleshooting Tip
If you're still getting the error, double-check your CSV's header row! Maybe the column is named Timestamp (capital T) instead of timestamp, or there's a typo. You can quickly verify the header with:
print(pd.read_csv(os.path.join(self.csv_dir, '%s.csv' % s), nrows=0).columns)
内容的提问来源于stack exchange,提问作者Bazman

