如何基于多个CSV时间序列文件构建指定顺序的MultiIndex DataFrame
Let's walk through how to assemble your OHLC data exactly as you requested—sorted by DateTime first, followed by the ticker order ['ANZ', 'NAB', 'WBC'], and finally with your desired column sequence.
Step 1: Read Data with Ticker Identification
First, we'll tweak your original reading code to add a ticker column to each DataFrame (so we can track which stock each row belongs to). Also, note that pd.DataFrame.from_csv is deprecated, so we'll use pd.read_csv instead with explicit index handling:
import pandas as pd asxList = ['ANZ', 'NAB', 'WBC'] all_data = [] for asxCode in asxList: # Read CSV, set DateTime as index (adjust index_col if your date column has a different name) ohlcData = pd.read_csv(f"{asxCode}.CSV", header=0, index_col='DateTime', parse_dates=True) # Add a column to mark which ticker this data belongs to ohlcData['Ticker'] = asxCode # Append to our collection of DataFrames all_data.append(ohlcData)
Step 2: Combine and Sort Data
Next, we'll concatenate all DataFrames, then sort by two levels: first the DateTime index, then the ticker in your specified order. To enforce the custom ticker order (instead of default alphabetical sorting), we'll convert the Ticker column to a categorical type with your predefined sequence:
# Merge all individual DataFrames into one combined_df = pd.concat(all_data) # Convert Ticker to categorical to lock in your desired order combined_df['Ticker'] = pd.Categorical(combined_df['Ticker'], categories=asxList, ordered=True) # Sort first by DateTime index, then by Ticker (using our custom order) sorted_df = combined_df.sort_values(by=['DateTime', 'Ticker'])
Step 3: Reorder Columns
Finally, specify your desired column sequence and rearrange the DataFrame. For example, if you want columns in ['Ticker', 'Open', 'High', 'Low', 'Close', 'Volume'] order:
# Define your preferred column order (update to match your actual column names) desired_columns = ['Ticker', 'Open', 'High', 'Low', 'Close', 'Volume'] final_df = sorted_df[desired_columns]
Key Notes
- Ensure your CSV files have a DateTime column (adjust
index_colinread_csvif your date column uses a different name like 'Date'). - Using
pd.Categoricalis critical here—it guarantees the ticker sort follows['ANZ', 'NAB', 'WBC']instead of the default alphabetical order (which would be ANZ, WBC, NAB). - If your original OHLC columns have unique names, update
desired_columnsto match your actual dataset.
内容的提问来源于stack exchange,提问作者artDeco

