两段Python爬虫代码中while循环异常表现的原因咨询
两段Python爬虫while循环异常表现的原因解析
运行环境:Python 3.11、HomeBrew、macOS 13.1 x64
关联的fetch函数完整代码
import datetime import time import requests import bs4 import sqlite3 def wget(UA,url,tid): try: cache=requests.get(url,headers=UA,timeout=6.05) cache.raise_for_status() cache.encoding=cache.apparent_encoding return cache.text except: print("[DEBUG]:Failed to fetch tid %d."%tid) return "" def modify(cache,tid): try: soup=bs4.BeautifulSoup(cache,"html.parser") result=soup.prettify() return result except: print("[DEBUG]:Failed to modify tid %d."%tid) return "" def export(page,tid): database=sqlite3.connect("post-cache1.sqlite3") c=database.cursor() c.execute("CREATE TABLE IF NOT EXISTS posts(tid INT NOT NULL PRIMARY KEY, content MEDIUMTEXT NOT NULL)") try: c.execute("INSERT INTO posts VALUES (?,?)",(tid,page)) database.commit() database.close() except: print("[DEBUG]:Failed in executing database for tid %d."%tid) database.close() def fetch(tid): UA={"User-Agent":"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.1 Safari/605.1.15"} url="https://****/forum.php?mod=viewthread&tid=%d"%tid print(url) origin=wget(UA,url,tid) mod=modify(origin,tid) export(mod,tid)
一、第一段代码:看似死循环却能稳定运行的原因
代码片段
def main(): interval=datetime.timedelta(microseconds=250) for i in range(1200000,1400000): #for i in range(1300000,1300010): start=datetime.datetime.now() end=start+interval print(start) fetch(i) realend=datetime.datetime.now() while realend<=end: time.sleep(end-realend) i+=1 main()
核心原因:while循环绝大多数情况下并未执行
你认为这是死循环,是因为realend在循环内没有更新,理论上只要进入循环就会一直满足realend <= end的条件,但实际运行中这个循环几乎不会被触发:
- 你的目标是每秒4次请求,即单次循环周期0.25秒,但
fetch(i)包含的网络请求、HTML解析、数据库操作,实际耗时很可能已经接近或超过0.25秒。 - 当
fetch(i)执行完毕后,realend = datetime.datetime.now()的时间已经晚于end(start + 0.25秒),此时realend <= end的条件不成立,while循环直接跳过,根本不会进入。 - 只有当
fetch(i)的耗时远小于0.25秒时,才会进入循环并陷入死循环——只是你的场景中从未出现这种情况,所以整体看起来能稳定以每秒4次的速率运行。
另外补充:代码中time.sleep(end-realend)存在语法问题,time.sleep()只接受数值类型(秒),而end-realend是datetime.timedelta对象,正常运行会抛出TypeError。你能运行可能是实际代码中做了转换(比如.total_seconds()),只是粘贴时遗漏了。
二、第二段代码:循环耗时远超预期的原因
代码片段
def main(): interval=datetime.timedelta(microseconds=250) for i in range(1200000,1400000): #for i in range(1300000,1300010): start=datetime.datetime.now() end=start+interval print(start) fetch(i) while datetime.datetime.now()<=end: time.sleep(0.000001) i+=1 main()
核心原因:fetch(i)的实际耗时远超0.25秒,循环周期由fetch决定
你期望循环周期稳定在0.25秒,但实际耗时0.5~1秒,速率降至每秒1次,问题出在fetch(i)的三个关键步骤:
- 网络请求限速:目标网站大概率存在反爬策略,当你尝试以每秒4次的频率请求时,网站会故意延迟响应、限制请求速率,导致
requests.get()的实际等待时间大幅增加(远超0.25秒)。 - 数据库操作低效:
export()函数每次执行都重新连接SQLite数据库、检查表是否存在,这些磁盘IO操作本身就有一定耗时,多次重复执行会累积延迟。 - 忙等待的副作用:while循环中使用
datetime.datetime.now()反复判断+1微秒休眠的方式属于忙等待,会持续占用CPU资源,可能导致Python解释器处理网络响应、数据库操作时的调度优先级降低,间接增加整体耗时。
简单来说,fetch(i)的实际执行时间已经远远超过了你设定的0.25秒间隔,所以while循环根本没有执行的机会,整个循环的耗时完全由fetch(i)的耗时决定,最终表现为每秒1次左右的请求速率。
内容的提问来源于stack exchange,提问作者Merry
相关产品推荐
相关产品推荐

