如何从Pandas DataFrame生成可用于read_csv的dtype字典?
The core issue here is that your Lookup DataFrame stores type names as strings (like "np.uint16"), but pd.read_csv expects actual type objects (like np.uint16) or recognized string aliases for those types. Here are a few safe, efficient ways to convert those strings into usable types:
Option 1: Use Pandas-Recognized String Aliases (Simplest)
Pandas understands numpy dtype string aliases (e.g., "uint16" instead of "np.uint16"). You can strip the "np." prefix from your type strings to make them compatible:
import pandas as pd import numpy as np # Modify the Type column to remove "np." prefix Lookup['Type'] = Lookup['Type'].str.replace('np.', '') # Create the dtype dictionary dtype_dict = dict(zip(Lookup['Variable'], Lookup['Type'])) # Now use it in read_csv df = pd.read_csv("input.csv", dtype=dtype_dict)
This works because pandas maps strings like "uint16" directly to np.uint16 under the hood, and "object" is already a valid alias for the Python object type.
Option 2: Create an Explicit Type Mapping (Safest)
If you want full control over which types are allowed, create a dictionary that maps your string type names to actual type objects. This avoids any risk of unintended code execution (unlike eval):
import pandas as pd import numpy as np # Define your type mapping type_mapping = { "object": object, "np.uint16": np.uint16, # Add other types from your Lookup DataFrame here (e.g., "np.int32": np.int32) } # Map the string types to actual objects Lookup['Type'] = Lookup['Type'].map(type_mapping) # Create the dtype dictionary dtype_dict = dict(zip(Lookup['Variable'], Lookup['Type'])) # Use in read_csv df = pd.read_csv("input.csv", dtype=dtype_dict)
This is ideal if you have a fixed set of types and want to avoid any potential issues with eval.
Option 3: Use eval() (Flexible but Caution Needed)
If you have many different numpy types and don't want to manually list them all, you can use eval() to convert the string type names into actual type objects. Only use this if you trust the content of your Lookup DataFrame (since eval executes arbitrary code):
import pandas as pd import numpy as np # Convert string type names to actual type objects using eval Lookup['Type'] = Lookup['Type'].apply(eval) # Create the dtype dictionary dtype_dict = dict(zip(Lookup['Variable'], Lookup['Type'])) # Use in read_csv df = pd.read_csv("input.csv", dtype=dtype_dict)
This works because eval("np.uint16") returns the actual np.uint16 type object, which pandas can use in read_csv.
Bonus: Verify the Dictionary
Before using the dtype dictionary, you can check that all values are valid type objects:
print({k: type(v) for k, v in dtype_dict.items()}) # Should output something like {'Var1': <class 'type'>, 'Var2': <class 'numpy.dtype'>}
All these methods will generate a valid dtype dictionary that pd.read_csv can use to load your large CSV efficiently, reducing memory usage by applying the correct data types upfront.
内容的提问来源于stack exchange,提问作者Richard Kapustynskyj

