Playwright异步函数Page.inner_html未await警告的排查求助
问题描述
我使用Python 3.11编写了一个Playwright异步函数,用于从实验室信息系统提取检测结果。尽管已为所有可见的Playwright操作添加await,但仍出现RuntimeWarning,提示coroutine 'Page.inner_html' was never awaited。代码可正常运行,但该警告始终存在,希望能找到问题所在并解决。
报错信息
RuntimeWarning: coroutine 'Page.inner_html' was never awaited if j >= 10 : RuntimeWarning: Enable tracemalloc to get the object allocation traceback
functions.py代码片段
... async def scrape_results(page, links_master): master_results=pd.DataFrame() j=1 for link in links_master: await page.goto("https://trakcarelabwebview.nhls.ac.za/trakcarelab/csp/system.Home.cls" + link) await page.wait_for_load_state(state="networkidle") await page.wait_for_selector(loading_icon, state="hidden") await page.wait_for_timeout(500) await page.wait_for_selector(loading_icon, state="hidden") await page.locator(test_item_caption).text_content(timeout=3000) == "TestItem" #Getting results off the page html = await page.content() soup = BeautifulSoup(html, "lxml") tables = soup.find_all('table') dfs = pd.read_html(str(tables)) df=dfs[1] try: word_result = BeautifulSoup(page.inner_html("#web_EPVisitTestSet_WordResult_0-ngForm > div > div:nth-child(2)"), "lxml").text_content() except: episode = await page.locator("#web_EPVisitNumber_List_Banner-row-0-item-Episode").text_content() collectiondate = await page.locator("#web_EPVisitTestSet_Result_0-item-CollectionDate").text_content() collectiontime = await page.locator("#web_EPVisitTestSet_Result_0-item-CollectionTime").text_content() resultdate = await page.locator("#web_EPVisitTestSet_Result_0-item-ResultDate").text_content() resulttime = await page.locator("#web_EPVisitTestSet_Result_0-item-ResultTime").text_content() specimen_info = { 'episode': episode, 'collectiondate': collectiondate, 'collectiontime': collectiontime, 'resultdate':resultdate, 'resulttime':resulttime } cleanup_results(df) #df = df.assign("Episode"==specimen_info.episode) #df = df.assign("Episode"==specimen_info['episode']) df=df.assign(Episode=specimen_info['episode']) df=df.assign(Collection_Date=specimen_info["collectiondate"]) df=df.assign(Collection_Time=specimen_info["collectiontime"]) df=df.assign(Result_Date=specimen_info["resultdate"]) df=df.assign(Result_Time=specimen_info["resulttime"]) master_results=pd.concat([master_results,df], ignore_index=True) j+=1 print("Result number __" + str(j) + "__ extracted") if j >= 10 : print(master_results) try: print(word_result) break except: print("No word_result found...") break ...
调用函数的代码
import asyncio from playwright.async_api import Playwright, async_playwright, expect import functions #The elements to interact with username_box = "#SSUser_Logon_0-item-USERNAME" password_box = "#SSUser_Logon_0-item-PASSWORD" ... _username = "User" _password = "Password" _param = "MRN139822539" async def run(playwright: Playwright, _username, _password, _param) -> None: browser = await playwright.chromium.launch(headless=False, slow_mo=50) try: context = await browser.new_context(storage_state="state.json") print("Storage state loaded...") except: print("Error upon loading storage state - continuing with login...") page = await context.new_page() await page.goto("https://the-web-site-im-visiting.com") all_links = await functions.get_links(page) await functions.scrape_results(page, all_links) context = await browser.new_context(storage_state="state.json") await page.wait_for_timeout(2000) await context.close() await browser.close() async def main() -> None: async with async_playwright() as playwright: await run(playwright, _username, _password, _param) asyncio.run(main())
解决方案
核心问题定位
警告的直接原因是**page.inner_html调用未添加await**。Playwright异步API中,page.inner_html是协程函数,必须通过await等待其执行完成,否则会生成未被处理的协程对象,触发RuntimeWarning。
代码修改
将try块中的page.inner_html调用添加await:
try: # 先await获取innerHTML内容,再传入BeautifulSoup inner_html_content = await page.inner_html("#web_EPVisitTestSet_WordResult_0-ngForm > div > div:nth-child(2)") word_result = BeautifulSoup(inner_html_content, "lxml").text_content() except: # 原异常处理逻辑不变 ...
额外优化建议
代码中还有一行无效的判断逻辑:
await page.locator(test_item_caption).text_content(timeout=3000) == "TestItem"
这行代码仅等待元素文本加载,但未实际执行判断。如果需要验证文本内容,应修改为:
test_item_text = await page.locator(test_item_caption).text_content(timeout=3000) if test_item_text != "TestItem": # 根据需求添加处理逻辑,比如跳过当前循环或抛出异常 continue
内容的提问来源于stack exchange,提问作者NanoNerd
相关产品推荐
相关产品推荐

