为何指定不同文件名时data.table::fread读取文件耗时差异大?
fread is much slower with file:/// paths vs. direct local paths Great question! Let's break down the significant runtime difference you've observed when using data.table::fread() with two path formats for a local CSV file.
Your Test Setup & Results
You tested two approaches to read the same local file:
- Using the
file:///prefix:fread("file:///C:/Users/Desktop/ads.csv") - Using a direct local path:
fread("C:/Users/Desktop/ads.csv")
With microbenchmark, you saw clear evidence that the file:/// version is far slower:
microbenchmark( fread("file:///C:/Users/Desktop/ads.csv"), fread("C:/Users/Desktop/ads.csv") )
Unit: microseconds expr min lq mean median uq max neval cld fread("file:///C:/Users/Desktop/ads.csv") 5755.975 6027.4735 6696.7807 6235.3365 6506.652 41257.476 100 b fread("C:/Users/Desktop/ads.csv") ...
The Root Cause of the Speed Difference
The gap comes down to how fread handles the two path types under the hood:
Direct local paths use optimized native file access
When you pass a standard local path (likeC:/Users/Desktop/ads.csv),freadleverages your operating system's native, low-level file reading APIs (e.g.,CreateFileon Windows,openon Unix-like systems). These APIs are built specifically for fast, direct access to local file systems—they minimize overhead and letfreadpull data from the file as efficiently as possible, which is part of whyfreadis known for its speed with local datasets.file:///paths trigger generalized URL-handling logic
Adding thefile:///prefix tellsfreadto treat the path as a URL instead of a local file path. This switches it to a completely different code path designed for remote or network-based resources (like HTTP/FTP URLs):- It first parses the URL using dedicated URL-processing libraries, adding extra steps that aren't needed for a simple local path.
- It uses a generic URL-based file reader (often relying on libraries like libcurl) instead of the native local file APIs. These generic readers include layers of abstraction to support multiple URL schemes, which adds significant overhead even when accessing a local file.
- There's extra work involved in converting the URL path back to a valid local file path, plus additional checks and buffer management that aren't required for direct local access.
In short, the file:/// prefix forces fread to use a one-size-fits-all URL-handling pipeline that's overkill for local files, while the direct path uses the optimized, local-file-specific code that makes fread so performant.
内容的提问来源于stack exchange,提问作者Ash Reddy

