为何无法使用await调用异步函数scrape_season?
解决Python异步爬取中"await allowed only within async function"错误及代码优化
错误根源
你碰到的"await" allowed only within async function错误,本质是**await关键字只能在被async修饰的函数内部使用**,而你直接在顶层的for循环里调用了await scrape_season(season),违反了Python异步语法规则。
核心修复方案
- 把遍历赛季的主逻辑封装到一个异步函数里(比如命名为
main) - 使用
asyncio.run()启动这个主异步函数,这是Python 3.7+启动异步程序的标准方式
附带代码优化(解决隐藏问题)
除了修复语法错误,你的代码还有几个影响功能和效率的问题,一并修正:
- 替换同步sleep为异步sleep:原代码用
time.sleep会阻塞整个异步事件循环,换成asyncio.sleep才能让异步任务并行处理 - 修正缩进错误:
scrape_season里的HTML获取和保存代码没缩进在for循环里,导致只会处理最后一个页面,现在调整缩进让每个页面都被处理 - 复用Playwright浏览器实例:原代码每次请求都重新启动浏览器,效率极低,现在改成启动一次浏览器复用,大幅提升爬取速度
- 提前创建目录:避免保存文件时因为目录不存在报错
修正后的完整代码
import os import asyncio from bs4 import BeautifulSoup from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeout SEASONS = list(range(2016,2023)) DATA_DIR = 'data' STANDINGS_DIR = os.path.join(DATA_DIR, 'standings') SCORES_DIR = os.path.join(DATA_DIR, 'scores') # 提前创建所需目录,避免保存时出错 os.makedirs(STANDINGS_DIR, exist_ok=True) os.makedirs(SCORES_DIR, exist_ok=True) async def get_html(page, url, selector, sleep=5, retries=3): html = None for i in range(1, retries+1): await asyncio.sleep(sleep * i) try: await page.goto(url) print(await page.title()) html = await page.inner_html(selector) except PlaywrightTimeout: print(f'Timeout error on {url}') continue else: break return html async def scrape_season(page, season): url = f'https://www.basketball-reference.com/leagues/NBA_{season}_games.html' html = await get_html(page, url, '#content .filter') soup = BeautifulSoup(html, 'html.parser') links = soup.find_all('a') hrefs = [l['href'] for l in links] standings_pages = [f"https://basketball-reference.com{l}" for l in hrefs] for url in standings_pages: save_path = os.path.join(STANDINGS_DIR, url.split("/")[-1]) if os.path.exists(save_path): print(f"已存在,跳过:{save_path}") continue html = await get_html(page, url, '#all_schedule') if html: with open(save_path, 'w+', encoding='utf-8') as f: f.write(html) print(f"已保存:{save_path}") async def main(): async with async_playwright() as p: # 启动浏览器(headless=True可以后台运行,不需要显示窗口) browser = await p.firefox.launch(headless=True) page = await browser.new_page() # 遍历赛季爬取 for season in SEASONS: await scrape_season(page, season) await browser.close() if __name__ == "__main__": asyncio.run(main())
关键改动说明
- 主异步函数
main():作为程序入口,负责启动浏览器、调度爬取任务、关闭浏览器 - 复用浏览器实例:只启动一次浏览器,把page对象传给其他函数,减少资源开销
- 异步sleep:确保事件循环在等待时能处理其他任务,提升异步效率
- 缩进修正:把文件保存逻辑放到for循环内部,保证每个赛程页面都被保存
- 编码与目录处理:指定utf-8编码避免乱码,提前创建目录防止路径错误
内容的提问来源于stack exchange,提问作者wade watts
相关产品推荐
相关产品推荐

