如何在pandas的read_csv()方法中指定有序分类类型?
Great question! Unfortunately, pandas doesn't support directly passing a string like 'ordered category' to the dtype parameter of read_csv()—the 'category' string alias only creates unordered categorical columns. But there's a clean way to achieve exactly what you want: create a pd.CategoricalDtype instance with ordered=True upfront, then pass that to dtype.
Step-by-Step Solution
Since you mentioned you can automatically determine the sort order (either as strings or numbers), here's how to set up an ordered category during the initial read_csv() call:
Define the ordered category type
First, generate the sorted list of categories. If you don't know the values in advance, do a quick "preview" read of just the target column to extract values, sort them, and create the custom dtype.Pass the custom dtype to
read_csv()
Use thepd.CategoricalDtypeinstance in thedtypedictionary instead of a plain string.
Example Code (Matching Your Scenario)
Let's adapt your example to make the BAR column an ordered category sorted by numeric value:
#!/usr/bin/env python3 import io import pandas as pd # Your original CSV data csv_data = """FOO;BAR\n 1;20204\n 5;20183\n 5;20182\n 4;20212\n""" # Step 1: Get sorted categories (sorted numerically here) # Temporary read to extract just the BAR column values temp_df = pd.read_csv( io.StringIO(csv_data), usecols=['BAR'], sep=';', keep_default_na=False ) # Convert to integers to sort numerically, then back to strings for categories sorted_categories = temp_df['BAR'].astype(int).sort_values().astype(str).unique() # Create the ordered categorical dtype bar_ordered_dtype = pd.CategoricalDtype( categories=sorted_categories, ordered=True ) # Step 2: Read full data with the ordered dtype df = pd.read_csv( io.StringIO(csv_data), keep_default_na=False, header=0, sep=';', dtype={'BAR': bar_ordered_dtype} ) # Verify the result print("BAR column (ordered category):") print(df.BAR) print("\nIs this an ordered category?", df.BAR.cat.ordered) # Outputs True print("\nCategory order:", df.BAR.cat.categories)
Key Notes
- If you need lexicographical string sorting instead of numeric, skip the
astype(int)step:sorted_categories = sorted(temp_df['BAR'].unique()) - If you already know the full list of categories and their order upfront, you can define
sorted_categoriesdirectly (e.g.,['20182', '20183', '20204', '20212']) without the temporary read. - This approach ensures the column is an ordered category during the initial read, which avoids any post-processing steps as you requested.
内容的提问来源于stack exchange,提问作者buhtz

