Python脚本在StocksList.txt含多行内容时出现挂起问题求助
解决Python脚本处理多行股票列表时的挂起问题
问题根源分析
你的脚本出现挂起的核心原因有三个:
- 未启用超时控制:代码中定义了
timeout变量,但requests.get调用时未传入该参数,导致请求可能无限等待服务器响应。 - 无请求间隔触发反爬:连续高频请求目标网站,服务器可能会限制或阻塞你的请求,导致程序停滞。
- 异常处理过于宽泛:外层
except Exception捕获所有异常,无法定位具体问题(比如网络错误、JSON解析失败等),也无法及时终止异常请求。
修复后的完整代码
import requests import time import chime infile = 'C:/Volume/PythonProjects/Final/Marketsmithindia/StocksList.txt' outfile = 'C:/Volume/PythonProjects/Final/Marketsmithindia/InstrumentID.csv' failfile = 'C:/Volume/PythonProjects/Final/Marketsmithindia/Failout.txt' timeout = 10 search_url = "https://msi-gcloud-prod.appspot.com/gateway/simple-api/ms-india/instr/srch.json" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/103.0.0.0 Safari/537.36" } # 计算运行时间 start = time.time() with open(infile, 'r', encoding='utf-8') as IN, \ open(outfile, 'w', encoding='utf-8') as OUT, \ open(failfile, 'w', encoding='utf-8') as FAIL: for idx, DATA in enumerate(IN, 1): STOCK = DATA.strip() if not STOCK: # 跳过空行,避免无效请求 continue try: params = {"text": STOCK} # 启用超时设置,避免请求无限等待 response = requests.get(search_url, params=params, headers=headers, timeout=timeout) response.raise_for_status() # 主动触发HTTP错误异常 data = response.json() # 检查返回数据结构完整性,避免索引越界或键缺失 if "response" not in data or "results" not in data["response"] or len(data["response"]["results"]) == 0: print(f"{STOCK}: 无匹配结果", file=FAIL) continue InstrumentID = data["response"]["results"][0]["instrumentId"] line = f"{STOCK},{InstrumentID}" print(line, file=OUT) # 添加请求间隔,降低被反爬的概率 if idx % 5 == 0: time.sleep(1) # 每5次请求停顿1秒 else: time.sleep(0.2) # 单次请求后短停顿 except requests.exceptions.RequestException as e: print(f"{STOCK}: 请求错误 - {str(e)}", file=FAIL) except KeyError as e: print(f"{STOCK}: 数据结构错误 - 缺失键 {str(e)}", file=FAIL) except Exception as e: print(f"{STOCK}: 未知错误 - {str(e)}", file=FAIL) end = time.time() with open(failfile, "a", encoding='utf-8') as file1: runtime = round((end - start), 3) print(f'总运行时间: {runtime} 秒', file=file1) chime.theme('material') chime.success()
关键修改说明
- 强制超时控制:在
requests.get中加入timeout=timeout,确保请求超过10秒自动终止,彻底解决挂起问题。 - 请求间隔优化:通过
time.sleep()添加请求间隔,避免高频请求触发服务器限制,平衡效率与合规性。 - 精细化异常处理:拆分不同类型的异常(请求错误、键错误、未知错误),便于后续排查具体问题;同时主动检查返回数据结构,避免索引越界。
- 空行过滤:跳过文件中的空行,减少无效请求次数。
- 编码指定:打开文件时明确指定
encoding='utf-8',避免因编码不一致导致的读取/写入错误。
内容的提问来源于stack exchange,提问作者erukumk
相关产品推荐
相关产品推荐

