使用Dataset.map()时函数哈希失败警告的原因及解决方法
在处理英法翻译JSON数据集(示例数据如下)时,我编写了分词前的预处理函数,调用后出现了哈希相关的警告。目前训练流程能正常运行,但想搞清楚警告的原因,以及如何消除它。
示例数据集:
{"translation": {"en": "Hello", "fr": "Bonjour"}} {"translation": {"en": "Enjoy your meal", "fr": "bon appétit"}} {"translation": {"en": "Please", "fr": "s'il te plaît"}}
预处理函数代码:
source_lang = "en" target_lang = "fr" def preprocess_dataset(examples): inputs = [example[source_lang] for example in examples["translation"]] targets = [example[target_lang] for example in examples["translation"]] model_inputs = tokenizer(inputs, return_tensors="pt", text_target=targets, padding="max_length", max_length=200, truncation=True) return model_inputs
调用代码:
tokenized_dataset = dataset.map(preprocess_dataset, batched=True)
运行后弹出的警告:
Parameter 'function'=<function preprocess_dataset at 0x7f968c4d4790> of the transform datasets.arrow_dataset.Dataset._map_single couldn't be hashed properly, a random hash was used instead. Make sure your transforms and parameters are serializable with pickle or dill for the dataset fingerprinting and caching to work. If you reuse this transform, the caching mechanism will consider it to be different from the previous calls and recompute everything. This warning is only showed once. Subsequent hashing failures won't be showed.
这个警告的核心是dataset.map()需要对传入的预处理函数做哈希计算——目的是生成数据集的唯一指纹,实现缓存复用:如果两次预处理的函数逻辑、参数完全一致,就直接调用之前缓存的结果,不用重复计算。
但你的preprocess_dataset函数依赖了外部定义的source_lang、target_lang和tokenizer变量,这些外部引用无法被pickle或dill序列化,导致函数的哈希计算失败,只能用随机哈希替代。这会导致后续重复调用该函数时,缓存机制会判定这是全新的变换操作,不得不重新执行预处理,浪费计算资源。
方法一:将外部变量转为函数参数
把source_lang、target_lang和tokenizer作为参数传入函数,让函数不依赖外部环境,就能正常被序列化:
def preprocess_dataset(examples, tokenizer, source_lang="en", target_lang="fr"): inputs = [example[source_lang] for example in examples["translation"]] targets = [example[target_lang] for example in examples["translation"]] model_inputs = tokenizer(inputs, return_tensors="pt", text_target=targets, padding="max_length", max_length=200, truncation=True) return model_inputs
调用时通过fn_kwargs传递参数:
tokenized_dataset = dataset.map(preprocess_dataset, batched=True, fn_kwargs={"tokenizer": tokenizer})
方法二:用类封装预处理逻辑
把预处理逻辑封装到类中,实现__call__方法,类的实例可以被正确序列化:
class Preprocessor: def __init__(self, tokenizer, source_lang="en", target_lang="fr"): self.tokenizer = tokenizer self.source_lang = source_lang self.target_lang = target_lang def __call__(self, examples): inputs = [example[self.source_lang] for example in examples["translation"]] targets = [example[self.target_lang] for example in examples["translation"]] model_inputs = self.tokenizer(inputs, return_tensors="pt", text_target=targets, padding="max_length", max_length=200, truncation=True) return model_inputs
调用时先实例化再传入:
preprocessor = Preprocessor(tokenizer) tokenized_dataset = dataset.map(preprocessor, batched=True)
方法三:禁用缓存检查(不推荐)
如果不在乎缓存失效、重复计算的问题,可以通过参数跳过哈希检查,但这会导致每次调用都重新执行预处理,仅适合临时测试:
tokenized_dataset = dataset.map(preprocess_dataset, batched=True, load_from_cache_file=False)
内容的提问来源于stack exchange,提问作者Raptor

