请求协助排查导入71MB Excel文件时的IndexError: pop from empty stack问题
Hey there, that IndexError: pop from empty stack when importing your 71MB Excel file is a common gotcha—usually tied to either library version quirks or wonky formatting in the file itself. Let's walk through the fixes that usually resolve this:
1. 先升级相关依赖库
This error pops up a lot in older versions of openpyxl (the default engine pandas uses for .xlsx files). Let's make sure you're running the latest stable versions:
pip install --upgrade pandas openpyxl
2. 切换Excel解析引擎
If upgrading doesn't work, try a different engine to parse the file. Each engine handles edge cases differently:
- Use xlrd (note: xlrd 2.0+ drops support for .xlsx, so install an older version if needed):
pip install xlrd==1.2.0import pandas as pd df = pd.read_excel("weather.xlsx", engine="xlrd") - Use pyxlsb (great for large .xlsx files since it reads the binary format):
pip install pyxlsbimport pandas as pd df = pd.read_excel("weather.xlsx", engine="pyxlsb")
3. 检查并修复Excel文件本身
Sometimes the file has hidden issues like corrupted cells, hidden sheets, or broken merged ranges:
- Open the file manually in Excel/LibreOffice—if you get a "repair file" prompt, let it repair, then save the fixed version and try importing again.
- Delete any hidden worksheets or overly complex merged cells, then re-save the file.
- If your data structure allows, save the file as CSV (File > Save As > CSV) and import that instead—CSV parsing is way more stable for large datasets:
df = pd.read_csv("weather.csv", low_memory=False)
4. 分块读取大文件
If memory constraints are causing parsing glitches, try reading the file in chunks to avoid overwhelming your system:
import pandas as pd # 每次读取10000行,可根据你的内存调整 chunk_iter = pd.read_excel("weather.xlsx", engine="openpyxl", chunksize=10000) df_list = [] for chunk in chunk_iter: df_list.append(chunk) # 合并所有块 final_df = pd.concat(df_list, ignore_index=True)
5. 重新下载文件
Double-check that your downloaded file isn't corrupted—sometimes partial downloads cause unexpected errors. Re-download the file, confirm it's ~71MB, then try importing again.
If none of these work, feel free to share a snippet of your import code and any additional error context, and we can dig deeper!
内容的提问来源于stack exchange,提问作者Vanna

