关于PyArrow ParquetWriter增量写入Parquet文件时未关闭前的可见性问题
Hey there! Great question—let me break this down for you based on how PyArrow's ParquetWriter works under the hood.
First off, your experimental observation is spot-on: the partially written Parquet file won't be visible or readable by standard Parquet clients until you call close(). This isn't just due to memory caching—it's rooted in how the Parquet format and PyArrow's implementation work together.
Here's the key detail: Parquet files rely on a final metadata block at the end of the file to define the structure of the data (like column schemas, data page locations, etc.). PyArrow's ParquetWriter holds onto this metadata until you explicitly call close()—it only writes that critical block once the writer is finalized. Without this metadata, any Parquet reader (whether it's PyArrow itself, parquet-tools, or another library) will either throw an error saying the file is invalid, or won't be able to parse any usable data from it.
While some data pages might get flushed to disk during incremental writes (depending on your OS's file cache and PyArrow's internal buffer settings), those are just fragmented pieces—nothing a valid Parquet reader can make sense of.
Since your main goal is to prevent other clients from accessing the incomplete file, here are a couple of tips to reinforce this:
- Test it directly: Mid-write, try opening the file with
pyarrow.parquet.ParquetFile()—you'll get an error about missing metadata or an invalid file structure. - Use a temporary file + atomic rename: For extra safety (especially in shared file systems), write to a temporary filename (like
my_data.parquet.tmp) first. Once you've calledclose()and confirmed the write is complete, rename it to your target filename. Most operating systems make file renaming an atomic operation, so other clients will only ever see the fully finished, valid Parquet file.
To circle back to your clarification: you don't have to worry about others seeing the incomplete file accidentally—Parquet's design and PyArrow's implementation already prevent that. The temp file trick is just a solid best practice to eliminate any edge cases.
备注:内容来源于stack exchange,提问作者Gabriele Giuseppini

