You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python异步网页抓取问题:run_forever()无法执行scrape函数

问题:scrape函数无法通过run_forever()实现每2秒永久运行

用户尝试让以下scrape函数每2秒永久执行:

import requests
from bs4 import BeautifulSoup
import asyncio

async def scrape():
    test = []
    r = requests.get(coin_desk)
    soup = BeautifulSoup(r.text, features='xml') 
    title = soup.find_all('title')[2]
    await asyncio.sleep(2)
    for x in title:
        test.append(x)
        print(test)

主函数代码:

try:
    loop = asyncio.new_event_loop()
    asyncio.set_event_loop(loop)
    loop.run_until_complete(scrape())
except (KeyboardInterrupt, SystemExit):
    pass

但将run_until_complete(scrape())替换为run_forever()后,程序仅永久运行却无任何输出,scrape函数未执行。


解决方案

核心问题

run_forever()仅负责启动事件循环,但不会自动调度任何协程任务;同时原scrape函数仅执行一次,无法实现重复运行的需求。

修改步骤

  1. 让scrape函数永久循环:在函数内添加while True循环,确保每次执行完成后等待2秒再重复执行
  2. 将协程任务加入事件循环:使用loop.create_task()把scrape协程提交到事件循环中,run_forever()才会调度执行该任务
  3. 修复未定义变量:补充coin_desk的URL定义,避免运行报错

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import asyncio

# 补充定义目标RSS源URL
coin_desk = "https://www.coindesk.com/feed"

async def scrape():
    while True:  # 添加永久循环逻辑
        test = []
        # 注意:requests是同步阻塞库,会阻塞事件循环,追求性能建议改用aiohttp
        r = requests.get(coin_desk)
        soup = BeautifulSoup(r.text, features='xml') 
        title = soup.find_all('title')[2]
        for x in title:
            test.append(x)
            print(test)
        await asyncio.sleep(2)  # 移到循环末尾,确保执行完一次再等待

try:
    loop = asyncio.new_event_loop()
    asyncio.set_event_loop(loop)
    # 将scrape协程作为任务提交到事件循环
    loop.create_task(scrape())
    loop.run_forever()
except (KeyboardInterrupt, SystemExit):
    # 优雅关闭事件循环
    loop.close()

额外优化建议

原代码中requests.get()是同步阻塞操作,会卡住整个事件循环。如果需要更高性能,建议替换为异步HTTP库aiohttp,示例如下:

import aiohttp
from bs4 import BeautifulSoup
import asyncio

coin_desk = "https://www.coindesk.com/feed"

async def scrape():
    async with aiohttp.ClientSession() as session:
        while True:
            test = []
            async with session.get(coin_desk) as r:
                text = await r.text()
                soup = BeautifulSoup(text, features='xml') 
                title = soup.find_all('title')[2]
                for x in title:
                    test.append(x)
                    print(test)
            await asyncio.sleep(2)

try:
    loop = asyncio.new_event_loop()
    asyncio.set_event_loop(loop)
    loop.create_task(scrape())
    loop.run_forever()
except (KeyboardInterrupt, SystemExit):
    loop.close()

内容的提问来源于stack exchange,提问作者Snuskbetty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 04:05:55