You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas快速获取大体积Excel文件的表头名称?

提速建议

针对读取大体积Excel表头过慢的问题,给你几个实用的提速方案:

  • 换用轻量型Excel读取库
    放弃openpyxl,改用pyexcelerate——它专门针对快速读取Excel结构做了优化,读取表头不需要解析整个文件内容。安装后用下面的代码:

    from pyexcelerate import Workbook
    wb = Workbook()
    wb.load("E:\DATA\dbo.xlsx")
    get_col = wb.get_sheet(0).get_row(1)  # 第一行是表头,索引从1开始
    print(get_col)
    

    这个方法通常能把耗时压缩到几秒内。

  • 直接解析XLSX的压缩包结构
    XLSX本质是ZIP压缩包,里面的工作表数据存在xl/worksheets/sheet1.xml(如果是第一个工作表)里。我们可以用Python内置的zipfile直接解压读取表头节点,完全绕过Excel解析库,速度最快:

    import zipfile
    from xml.etree import ElementTree as ET
    
    with zipfile.ZipFile("E:\DATA\dbo.xlsx", 'r') as zf:
        # 读取第一个工作表的xml文件,若有多个工作表需调整文件名
        with zf.open('xl/worksheets/sheet1.xml') as f:
            tree = ET.parse(f)
            root = tree.getroot()
            # 命名空间处理
            ns = {'ss': 'http://schemas.openxmlformats.org/spreadsheetml/2006/main'}
            # 获取第一行的所有单元格
            header_row = root.find('.//ss:row[@r="1"]', ns)
            get_col = [cell.find('.//ss:v', ns).text for cell in header_row.findall('.//ss:c', ns)]
    print(get_col)
    

    这个方法几乎是瞬时完成,因为只读取和解析表头对应的一小段XML内容。

  • 优化pandas的读取参数(效果有限但简单)
    如果坚持用pandas,可以尝试指定engine='xlrd'(注意:xlrd 2.0+不再支持XLSX,需要安装xlrd==1.2.0),同时加上usecols=None和dtype=object减少类型推断开销:

    import pandas as pd
    get_col = list(pd.read_excel("E:\DATA\dbo.xlsx", nrows=1, engine='xlrd', usecols=None, dtype=object).columns)
    print(get_col)
    

内容的提问来源于stack exchange,提问作者sridharnetha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 19:51:29