You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何无法使用await调用异步函数scrape_season?

解决Python异步爬取中"await allowed only within async function"错误及代码优化

错误根源

你碰到的"await" allowed only within async function错误,本质是**await关键字只能在被async修饰的函数内部使用**,而你直接在顶层的for循环里调用了await scrape_season(season),违反了Python异步语法规则。

核心修复方案

  1. 把遍历赛季的主逻辑封装到一个异步函数里(比如命名为main)
  2. 使用asyncio.run()启动这个主异步函数,这是Python 3.7+启动异步程序的标准方式

附带代码优化(解决隐藏问题)

除了修复语法错误,你的代码还有几个影响功能和效率的问题,一并修正:

  • 替换同步sleep为异步sleep:原代码用time.sleep会阻塞整个异步事件循环,换成asyncio.sleep才能让异步任务并行处理
  • 修正缩进错误:scrape_season里的HTML获取和保存代码没缩进在for循环里,导致只会处理最后一个页面,现在调整缩进让每个页面都被处理
  • 复用Playwright浏览器实例:原代码每次请求都重新启动浏览器,效率极低,现在改成启动一次浏览器复用,大幅提升爬取速度
  • 提前创建目录:避免保存文件时因为目录不存在报错

修正后的完整代码

import os
import asyncio
from bs4 import BeautifulSoup
from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeout

SEASONS = list(range(2016,2023))
DATA_DIR = 'data'
STANDINGS_DIR = os.path.join(DATA_DIR, 'standings')
SCORES_DIR = os.path.join(DATA_DIR, 'scores')

# 提前创建所需目录,避免保存时出错
os.makedirs(STANDINGS_DIR, exist_ok=True)
os.makedirs(SCORES_DIR, exist_ok=True)

async def get_html(page, url, selector, sleep=5, retries=3):
    html = None 
    for i in range(1, retries+1):
        await asyncio.sleep(sleep * i)

        try: 
            await page.goto(url)
            print(await page.title())
            html = await page.inner_html(selector)
    
        except PlaywrightTimeout:
            print(f'Timeout error on {url}')
            continue
        else: 
            break
    return html

async def scrape_season(page, season):
    url = f'https://www.basketball-reference.com/leagues/NBA_{season}_games.html'
    html = await get_html(page, url, '#content .filter')

    soup = BeautifulSoup(html, 'html.parser')
    links = soup.find_all('a')
    hrefs = [l['href'] for l in links]
    standings_pages = [f"https://basketball-reference.com{l}" for l in hrefs]

    for url in standings_pages:
        save_path = os.path.join(STANDINGS_DIR, url.split("/")[-1])
        if os.path.exists(save_path):
            print(f"已存在,跳过:{save_path}")
            continue

        html = await get_html(page, url, '#all_schedule')
        if html:
            with open(save_path, 'w+', encoding='utf-8') as f:
                f.write(html)
            print(f"已保存:{save_path}")

async def main():
    async with async_playwright() as p:
        # 启动浏览器(headless=True可以后台运行,不需要显示窗口)
        browser = await p.firefox.launch(headless=True)
        page = await browser.new_page()
        
        # 遍历赛季爬取
        for season in SEASONS:
            await scrape_season(page, season)
        
        await browser.close()

if __name__ == "__main__":
    asyncio.run(main())

关键改动说明

  • 主异步函数main():作为程序入口,负责启动浏览器、调度爬取任务、关闭浏览器
  • 复用浏览器实例:只启动一次浏览器,把page对象传给其他函数,减少资源开销
  • 异步sleep:确保事件循环在等待时能处理其他任务,提升异步效率
  • 缩进修正:把文件保存逻辑放到for循环内部,保证每个赛程页面都被保存
  • 编码与目录处理:指定utf-8编码避免乱码,提前创建目录防止路径错误

内容的提问来源于stack exchange,提问作者wade watts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 02:17:56