读取TFRecord时反复出现DataLossError截断记录问题排查求助
I keep running into a
DataLossError: truncated record at ____error when reading my TFRecord files. I've regenerated the TFRecord files multiple times, but the error still persists. Here's my TFRecord writing code:with tf.python_io.TFRecordWriter(file_name) as writer: for _, row in id_table.iterrows(): id = row['id'] example = _process_id(id) if example is not None: writer.write(example.SerializeToString())
Let's break down the possible causes for this truncated record error, even after regenerating your TFRecords:
Silent failures in the
_process_idfunction
Your_process_id(id)might be generating incomplete or invalidtf.train.Exampleobjects without returningNone. For example, if serialization hits an unexpected edge case (like a missing feature or malformed data) mid-process, you could end up writing a partial byte string to the file—this is a common culprit for truncation errors.
Try adding error handling and validation to catch these issues:with tf.python_io.TFRecordWriter(file_name) as writer: for _, row in id_table.iterrows(): id = row['id'] try: example = _process_id(id) if example is not None: serialized = example.SerializeToString() # Skip empty or suspiciously small serialized data if len(serialized) < 10: print(f"Warning: Suspiciously small example for id {id}") continue writer.write(serialized) except Exception as e: print(f"Failed to process id {id}: {str(e)}") continueUnderlying file system issues
Even if your code is correct, disk-level problems can corrupt the written files. This includes insufficient disk space, intermittent network glitches (if writing to a remote drive), or disk corruption.- Check that your target disk has enough free space.
- Try writing to a local disk instead of a remote storage system to rule out network-related truncation.
- While the
withstatement should handle closing the writer properly, adding an explicitwriter.flush()before the block exits can ensure all data is written to disk, especially if your script exits abruptly.
Malformed source data in
id_table
Some rows in yourid_tablemight have corrupted or unexpected data that breaks the serialization process, even if_process_iddoesn't returnNone. For example, a feature with a mismatched data type or an empty tensor that doesn't serialize correctly.
Add logging to inspect examples before writing: print the structure ofexamplefor a handful of rows, or write a small test to serialize and immediately deserialize an example to confirm it's valid.TensorFlow version mismatches
If you're using different TensorFlow versions to write and read the TFRecord files, subtle serialization/deserialization incompatibilities can lead to truncation errors. Make sure you're using the same major (and ideally minor) version for both operations.
内容的提问来源于stack exchange,提问作者Jack

