如何修改Python脚本:无需频繁请求API即可即时获取新Post Title
解决方案:平衡即时性与请求频率的几种实现方式
针对你想即时获取新帖子标题又不想过度请求API的需求,以下是几种实用方案:
1. 长轮询(Long Polling)推荐
如果目标API支持长轮询(即服务器会持有请求直到有新内容或超时),这是最优解——既不会频繁发送请求,又能在新内容发布时立即收到通知。
修改后的脚本示例:
import requests import time url = 'https://example.com/api' headers = {"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36"} previous_titles = [] def fetch_titles_long_poll(): global previous_titles while True: try: # 发送长轮询请求,设置合理的超时时间(比如30秒) resp = requests.get(url, headers=headers, timeout=30) data = resp.json() new_titles = [title['title'] for title in data['latest']['posts'] if title['title'] not in previous_titles] for title in new_titles: current_time = time.strftime('%Y-%m-%d %H:%M:%S') print(f"[{current_time}] {title}") previous_titles.append(title) except requests.exceptions.Timeout: # 超时后立即重新发起请求,避免等待 continue except Exception as e: print(f"请求出错: {e}") # 出错时暂停5秒再重试,防止频繁报错 time.sleep(5) if __name__ == "__main__": fetch_titles_long_poll()
说明:长轮询依赖API支持,若API没有实现该机制,这个方法不生效。
2. 动态调整轮询间隔
如果API不支持长轮询,可以采用「发现新内容后缩短间隔,无新内容时恢复长间隔」的策略,平衡即时性和请求量。
修改后的脚本示例:
import requests import time url = 'https://example.com/api' headers = {"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36"} previous_titles = [] # 基础间隔和短间隔 BASE_INTERVAL = 45 SHORT_INTERVAL = 6 # 连续无新内容时恢复长间隔的次数 SHORT_INTERVAL_RETRY = 3 def fetch_titles_dynamic(): global previous_titles short_retry_count = 0 while True: try: resp = requests.get(url, headers=headers) data = resp.json() new_titles = [title['title'] for title in data['latest']['posts'] if title['title'] not in previous_titles] if new_titles: for title in new_titles: current_time = time.strftime('%Y-%m-%d %H:%M:%S') print(f"[{current_time}] {title}") previous_titles.append(title) # 发现新内容,重置短间隔重试次数,下次用短间隔 short_retry_count = 0 sleep_interval = SHORT_INTERVAL else: short_retry_count += 1 # 连续多次无新内容,恢复长间隔 if short_retry_count >= SHORT_INTERVAL_RETRY: sleep_interval = BASE_INTERVAL else: sleep_interval = SHORT_INTERVAL time.sleep(sleep_interval) except Exception as e: print(f"请求出错: {e}") # 出错时用基础间隔重试 time.sleep(BASE_INTERVAL) if __name__ == "__main__": fetch_titles_dynamic()
说明:这个方案无需API支持,通过动态调整间隔,在有新内容时保持高频率检查,无新内容时降低请求频率,有效减少不必要的请求。
3. 利用API的更新标识(如ETag或Last-Modified)
如果API返回ETag或Last-Modified响应头,可以在后续请求中携带这些标识,服务器会在内容未变化时返回304 Not Modified,客户端无需解析响应内容,减少处理开销(虽然请求次数不变,但能节省带宽和本地计算资源)。
示例代码片段:
# 初始化ETag和Last-Modified etag = None last_modified = None def fetch_titles_with_etag(): global previous_titles, etag, last_modified while True: headers_with_cache = headers.copy() if etag: headers_with_cache['If-None-Match'] = etag if last_modified: headers_with_cache['If-Modified-Since'] = last_modified resp = requests.get(url, headers=headers_with_cache) if resp.status_code == 304: # 内容未变化,直接等待 time.sleep(45) continue # 更新缓存标识 etag = resp.headers.get('ETag') last_modified = resp.headers.get('Last-Modified') data = resp.json() # 后续处理和原脚本一致... # 省略原有的新标题判断和打印逻辑
内容的提问来源于stack exchange,提问作者Dontpanic
相关产品推荐
相关产品推荐

