You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用xlrd读取XLS文件报错:Unsupported format, or corrupt file

问题:读取XLS文件时触发xlrd格式不支持错误

我尝试将文件夹//DAILY POSTINGS - MAY23/中的所有XLS文件读取到Pandas DataFrame中,使用的代码如下:

daily_postings_folder = '//DAILY POSTINGS - MAY23/'

daily_postings_list = glob.glob(os.path.join(daily_postings_folder, "*.xls"))

daily_postings_df = pd.concat((pd.read_excel(g, engine='xlrd', header=None) for g in daily_postings_list), ignore_index=True)
daily_postings_df.head()

运行后出现如下错误:

Traceback (most recent call last):
  File "z:\Trade Reconciliation\Testing\Python_Testing\daily_trade_rec.v3(DEV).py", line 20, in <module>
    daily_postings_df = pd.concat((pd.read_excel(g, engine='xlrd', header=None) for g in daily_postings_list), ignore_index=True)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\util\_decorators.py", line 311, in wrapper
    return func(*args, **kwargs)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\core\reshape\concat.py", line 347, in concat
    op = _Concatenator(
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\core\reshape\concat.py", line 401, in __init__
    objs = list(objs)
  File "z:\Trade Reconciliation\Testing\Python_Testing\daily_trade_rec.v3(DEV).py", line 20, in <genexpr>
    daily_postings_df = pd.concat((pd.read_excel(g, engine='xlrd', header=None) for g in daily_postings_list), ignore_index=True)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\util\_decorators.py", line 311, in wrapper
    return func(*args, **kwargs)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\io\excel\_base.py", line 457, in read_excel
    io = ExcelFile(io, storage_options=storage_options, engine=engine)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\io\excel\_base.py", line 1419, in __init__
    self._reader = self._engines[engine](self._io, storage_options=storage_options)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\io\excel\_xlrd.py", line 25, in __init__
    super().__init__(filepath_or_buffer, storage_options=storage_options)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\io\excel\_base.py", line 518, in __init__
    self.book = self.load_workbook(self.handles.handle)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\pandas\io\excel\_xlrd.py", line 38, in load_workbook
    return open_workbook(file_contents=data)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\xlrd\__init__.py", line 172, in open_workbook
    bk = open_workbook_xls(
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\xlrd\book.py", line 79, in open_workbook_xls
    biff_version = bk.getbof(XL_WORKBOOK_GLOBALS)
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\xlrd\book.py", line 1284, in getbof
    bof_error('Expected BOF record; found %r' % self.mem[savpos:savpos+8])
  File "C:\Users\tac6328\Anaconda3\lib\site-packages\xlrd\book.py", line 1278, in bof_error
    raise XLRDError('Unsupported format, or corrupt file: ' + msg)
xlrd.biffh.XLRDError: Unsupported format, or corrupt file: Expected BOF record; found b'\xef\xbb\xbf<?xml'

确认所有文件均显示为XLS格式,不清楚报错原因,寻求解决方法。


解决方法

错误原因

报错信息里的b'\xef\xbb\xbf<?xml'说明文件实际是XML格式的表格文件(大概率是被改了后缀的XLSX文件),而xlrd 2.0及以上版本只支持老式二进制XLS文件,不再兼容XLSX/XML格式,因此触发格式不支持错误。

具体解决方案

  • 方案1:改用openpyxl引擎读取
    先安装openpyxl(未安装时执行):

    pip install openpyxl
    

    修改代码中的引擎参数为openpyxl,它支持XLSX及伪装成XLS的XML格式文件:

    daily_postings_df = pd.concat((pd.read_excel(g, engine='openpyxl', header=None) for g in daily_postings_list), ignore_index=True)
    
  • 方案2:修正文件后缀
    右键文件查看属性,确认实际文件类型。如果是XLSX格式,批量将文件后缀改为.xlsx,再用对应引擎读取,避免格式混淆。

  • 方案3:回退xlrd版本(不推荐)
    若必须使用xlrd,可安装1.2.0版本(支持XLSX),但该版本已停止维护,存在安全隐患:

    pip install xlrd==1.2.0
    

内容的提问来源于stack exchange,提问作者selereth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 16:47:49