如何根据HTTP请求传入信息为read_csv读取的列指定对应数据类型?
Hey there! Let's fix this issue step by step. The problem with your current code is that the nested loops are applying every type in type_fields to every column in list_fields—so by the end, all columns will end up with the last type in the list, which is definitely not what you want. We need to map each column name directly to its corresponding data type instead.
Option 1: Specify dtypes directly in read_csv (Recommended)
This is more efficient because it sets the data types while reading the CSV instead of converting after loading, which avoids potential parsing issues (like numbers being read as strings first).
First, we'll create a dictionary that maps each column name to its target type, then pass this to the dtype parameter of read_csv:
import pandas as pd project_id = request.data['project'] list_fields = request.POST.getlist('headers') type_fields = request.POST.getlist('type') # First, validate that we have a type for every column if len(list_fields) != len(type_fields): raise ValueError("Number of headers must match number of types") # Create a mapping of column names to their data types dtype_mapping = dict(zip(list_fields, type_fields)) # Read the CSV with the specified dtypes dataframe = pd.read_csv(file_path, header=0, dtype=dtype_mapping)
Option 2: Convert types after reading the CSV
If you need to do post-processing after loading (e.g., some columns need special handling), you can loop through the paired column-type pairs instead of using nested loops:
import pandas as pd project_id = request.data['project'] list_fields = request.POST.getlist('headers') type_fields = request.POST.getlist('type') # Validate length match first if len(list_fields) != len(type_fields): raise ValueError("Number of headers must match number of types") dataframe = pd.read_csv(file_path, header=0) # Iterate over paired column and type for field, tipo in zip(list_fields, type_fields): try: dataframe[field] = dataframe[field].astype(tipo) except ValueError as e: # Handle cases where conversion fails (e.g., non-numeric values for int type) print(f"Failed to convert column {field} to type {tipo}: {e}") # Optionally, fall back to a default type or skip dataframe[field] = dataframe[field].astype(str)
Key Tips
- Ensure the types in
type_fieldsare valid pandas/numpy data types—common options include'int64','float64','str','datetime64', or numpy-specific types likenp.int32. - Adding validation (like the length check) helps catch input mismatches early, making your code more robust.
- For date columns, use
read_csv'sparse_datesparameter instead ofastype, since dates often require specific parsing formats to convert correctly.
内容的提问来源于stack exchange,提问作者João Pedro Gomes

