在HTML中用py-script爬取网页遇BlockingIOError如何解决?
在PyScript中实现网页爬取遇到BlockingIOError错误的解决办法
问题代码
<body class="white-vertion black-bg"> <!-- Start Loader --> <p> <py-script> import ssl from urllib.request import urlopen from bs4 import BeautifulSoup context = ssl._create_unverified_context() result = urlopen("https://blog.naver.com/PostList.naver?blogId=woong3164&categoryNo=0&from=postList", context=context) bsObj = BeautifulSoup(result.read(), "html.parser") </py-script> </p>
错误信息
JsException(PythonError: Traceback (most recent call last): File "/lib/python3.10/urllib/request.py", line 1348, in do_open h.request(req.get_method(), req.selector, req.data, headers, File "/lib/python3.10/http/client.py", line 1282, in request self._send_request(method, url, body, headers, encode_chunked) File "/lib/python3.10/http/client.py", line 1328, in _send_request self.endheaders(body, encode_chunked=encode_chunked) File "/lib/python3.10/http/client.py", line 1277, in endheaders self._send_output(message_body, encode_chunked=encode_chunked) File "/lib/python3.10/http/client.py", line 1037, in _send_output self.send(msg) File "/lib/python3.10/http/client.py", line 975, in send self.connect() File "/lib/python3.10/http/client.py", line 1447, in connect super().connect() File "/lib/python3.10/http/client.py", line 941, in connect self.sock = self._create_connection( File "/lib/python3.10/socket.py", line 845, in create_connection raise err File "/lib/python3.10/socket.py", line 833, in create_connection sock.connect(sa) BlockingIOError: [Errno 26] Operation in progress During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/lib/python3.10/site-packages/_pyodide/_base.py", line 429, in eval_code .run(globals, locals) File "/lib/python3.10/site-packages/_pyodide/_base.py", line 300, in run coroutine = eval(self.code, globals, locals) File "", line 6, in File "/lib/python3.10/urllib/request.py", line 216, in urlopen return opener.open(url, data, timeout) File "/lib/python3.10/urllib/request.py", line 519, in open response = self._open(req, data) File "/lib/python3.10/urllib/request.py", line 536, in _open result = self._call_chain(self.handle_open, protocol, protocol + File "/lib/python3.10/urllib/request.py", line 496, in _call_chain result = func(*args) File "/lib/python3.10/urllib/request.py", line 1391, in https_open return self.do_open(http.client.HTTPSConnection, req, File "/lib/python3.10/urllib/request.py", line 1351, in do_open raise URLError(err) urllib.error.URLError: )'
错误原因
你误解了问题根源:这个错误不是SSL导致的,而是浏览器环境不允许同步阻塞式的网络请求。PyScript基于Pyodide运行在浏览器中,而浏览器的JavaScript环境强制要求网络操作必须是异步的,urllib.request.urlopen是同步阻塞方法,因此触发了BlockingIOError。
解决方法
使用Pyodide提供的异步网络请求工具pyodide.http.open_url,并配合async/await语法处理异步操作:
修改后的代码
<body class="white-vertion black-bg"> <!-- Start Loader --> <p> <py-script> import asyncio from pyodide.http import open_url from bs4 import BeautifulSoup async def fetch_page(): # 使用open_url发起异步请求,自动处理SSL和跨域问题 url = "https://blog.naver.com/PostList.naver?blogId=woong3164&categoryNo=0&from=postList" with open_url(url) as response: html_content = response.read() bsObj = BeautifulSoup(html_content, "html.parser") # 这里可以添加处理解析结果的代码,比如打印标题 print(bsObj.title.text) # 运行异步函数 asyncio.ensure_future(fetch_page()) </py-script> </p>
关键说明
pyodide.http.open_url是Pyodide专为浏览器环境设计的异步网络请求方法,兼容浏览器的CORS策略,无需手动处理SSL上下文。- 必须用
async def定义异步函数,并通过asyncio.ensure_future()来执行(PyScript顶层await需要额外配置,用此方法更稳妥)。 - 原代码中的
ssl._create_unverified_context()可以直接移除,open_url已处理HTTPS连接的SSL验证。
内容的提问来源于stack exchange,提问作者WOOOOONG_HERO
相关产品推荐
相关产品推荐

