使用pandas.read_excel读取URL中Excel文件时遇io.UnsupportedOperation: seek错误
Python 3.6从URL读取Excel文件解决
io.UnsupportedOperation: seek错误 问题场景
在Python 3.6中尝试从SharePoint URL读取Excel文件时,执行代码后触发io.UnsupportedOperation: seek错误。
原代码
import pandas as pd from urllib.request import Request, urlopen url = "https://<myOrg>.sharepoint.com/:x:/s/x-taulukot/Ec0R1y3l7sdGsP92csSO-mgBI8WCN153LfEMvzKMSg1Zzg?e=6NS5Qh" req = Request(url) req.add_header('User-Agent', 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:77.0) Gecko/20100101 Firefox/77.0') content = urlopen(req) pd.read_excel(content) print(df)
报错信息
(venv) miettinj@ramen:~/beta/python> python test.py Traceback (most recent call last): File "test.py", line 9, in <module> pd.read_excel(content) File "/srv/work/miettinj/beta/python/venv/lib/python3.6/site-packages/pandas/util/_decorators.py", line 296, in wrapper return func(*args, **kwargs) File "/srv/work/miettinj/beta/python/venv/lib/python3.6/site-packages/pandas/io/excel/_base.py", line 304, in read_excel io = ExcelFile(io, engine=engine) File "/srv/work/miettinj/beta/python/venv/lib/python3.6/site-packages/pandas/io/excel/_base.py", line 851, in __init__ if _is_ods_stream(path_or_buffer): File "/srv/work/miettinj/beta/python/venv/lib/python3.6/site-packages/pandas/io/excel/_base.py", line 800, in _is_ods_stream stream.seek(0) io.UnsupportedOperation: seek
解决方法
问题根源在于urlopen()返回的对象是不可seek的流,而pandas在判断文件类型时需要调用seek()方法,导致报错。
解决思路是将读取到的内容存入支持seek操作的BytesIO对象中,再传给pd.read_excel()。
修改后的代码
import pandas as pd from urllib.request import Request, urlopen from io import BytesIO # 导入BytesIO模块 url = "https://<myOrg>.sharepoint.com/:x:/s/x-taulukot/Ec0R1y3l7sdGsP92csSO-mgBI8WCN153LfEMvzKMSg1Zzg?e=6NS5Qh" req = Request(url) req.add_header('User-Agent', 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:77.0) Gecko/20100101 Firefox/77.0') content = urlopen(req).read() # 读取URL返回的全部二进制内容 excel_file = BytesIO(content) # 转为可seek的内存字节流 df = pd.read_excel(excel_file) # 正常读取Excel文件 print(df)
关键步骤说明
- 导入
BytesIO模块:用于在内存中创建可读写、支持seek操作的字节流容器。 - 调用
urlopen(req).read():一次性读取URL返回的全部二进制内容,避免流式读取的限制。 - 转为
BytesIO对象:将二进制内容存入内存流,满足pandas对可seek流的要求。 - 读取Excel:使用
pd.read_excel()读取内存流,完成文件解析。
内容的提问来源于stack exchange,提问作者Jaana
相关产品推荐
相关产品推荐

