相同硬件下Python3.6脚本在Win10运行极慢,Ubuntu却正常?
Let me walk through a recent task where I ran into a head-scratching performance discrepancy, and how I fixed it to work reliably across Windows and Ubuntu.
The Setup
I was tasked with merging three large CSV files into one, using their first shared attribute as the key. I wrote the initial code, tested it on my machine (same CPU as the client's, 16GB RAM, local SSD), and it finished in 10 seconds flat. But when the client ran it? It took a full 2 minutes.
After reworking the code and resubmitting, they mentioned this time it was running on an Ubuntu machine—so I had to make sure the fix worked seamlessly across both operating systems.
Why the Huge Performance Gap?
I dug into possible reasons, and these are the most likely culprits:
- File System & I/O Overhead: Even with SSDs, Windows' NTFS and Ubuntu's ext4 handle file operations differently. My initial code used small, repeated read calls which might have triggered more overhead on NTFS. Also, if the client was pulling files from a network drive (not a local SSD) that would absolutely kill speed—something I didn't account for in my local testing.
- Python Environment Mismatch: If the client was running an older Python version (like 3.7) without optimized libraries, that's a big one. For example, pandas' CSV parser got a massive speed boost in 3.10+ when using the
pyarrowengine, which I had enabled but they might not have installed or configured. - Background Processes: Windows real-time antivirus can scan every file read/write, adding significant latency that Ubuntu rarely has (unless they're running strict security tools in the background).
The Fixes That Worked Across Platforms
Here's what I changed to get consistent fast performance:
- Use PyArrow for CSV I/O: Swapped out the standard
csvmodule and basic pandas reads for the pyarrow engine, which is drastically faster for large file operations. Example code snippet:import pandas as pd # Read CSVs with pyarrow for optimized speed df1 = pd.read_csv("file1.csv", engine="pyarrow") df2 = pd.read_csv("file2.csv", engine="pyarrow") df3 = pd.read_csv("file3.csv", engine="pyarrow") # Merge on the shared key (assuming it's named 'record_id') merged_df = df1.merge(df2, on="record_id", how="outer").merge(df3, on="record_id", how="outer") # Write merged output with pyarrow merged_df.to_csv("merged_output.csv", engine="pyarrow", index=False) - Define Explicit Data Types: I stopped letting pandas guess data types, which saves processing time and prevents unnecessary memory bloat. For example:
dtype_map = {"record_id": str, "metric_a": float, "metric_b": int} df1 = pd.read_csv("file1.csv", engine="pyarrow", dtype=dtype_map) - Disable Unnecessary Processing: I turned off auto-date parsing (since none of the columns were date values) and skipped index creation when writing the final file, which cut down on extra, unneeded work.
- Optional Chunking for Edge Cases: For extra robustness, I added optional chunked merging logic in case the client's files were larger than expected—though in this case, the full dataset fit in RAM, but chunking ensures the code works even if memory is tight.
Final Thoughts
Never assume your local testing environment matches the client's! Always account for OS differences, library versions, and storage locations when optimizing code. Using optimized engines like pyarrow makes a huge difference for CSV operations across all platforms.
内容的提问来源于stack exchange,提问作者Dimitar Dimitrov

