使用Python从GitHub读取CSV文件时遇请求耗时过长问题求助
从GitHub Raw链接获取CSV时频繁出现长时间无响应的问题
我需要实时获取GitHub上某CSV文件的最新版本,对应的Raw链接为https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv。文件本身不大,正常1秒内就能完成获取,但尝试了两种导入到Pandas的方法后,都频繁出现长时间卡住无响应的情况(并非每次都会触发)。
尝试过的方法
方法一:直接用Pandas的read_csv
import pandas as pd url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv' df = pd.read_csv(url)
经常无报错但长时间运行,最终只能手动终止脚本。
方法二:先用requests获取再导入
import pandas as pd import requests import io url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv' response = requests.get(url) df = pd.read_csv(io.StringIO(response.content.decode('utf-8')))
卡在response = requests.get(url)这一行,耗时极长。
已知信息
- 网络连接正常
- 不能拉取仓库或下载到本地,必须实时获取最新版本
- Pandas版本:1.5.2
- Python版本:3.8.10
解决办法
1. 给请求添加超时时间
不管是requests还是pd.read_csv,都可以设置超时参数,避免无限等待:
- 针对requests:
response = requests.get(url, timeout=10) # 设置10秒超时,超时会抛出异常
- 针对Pandas的
read_csv(1.5+版本支持通过storage_options传递参数):
df = pd.read_csv(url, storage_options={'timeout': 10})
2. 增加重试机制
偶发无响应大概率是网络波动或GitHub CDN临时延迟,添加重试逻辑能解决大部分这类问题:
用tenacity库实现重试(需先安装pip install tenacity)
import pandas as pd import requests import io from tenacity import retry, stop_after_attempt, wait_exponential @retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10)) def get_csv_data(url): response = requests.get(url, timeout=10) response.raise_for_status() # 触发HTTP错误,比如404、500 return response.content.decode('utf-8') url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv' csv_content = get_csv_data(url) df = pd.read_csv(io.StringIO(csv_content))
自定义简单重试逻辑(无需额外安装库)
import pandas as pd import requests import io import time url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv' max_retries = 3 retry_delay = 2 for attempt in range(max_retries): try: response = requests.get(url, timeout=10) response.raise_for_status() df = pd.read_csv(io.StringIO(response.content.decode('utf-8'))) break except (requests.exceptions.RequestException, pd.errors.ParserError) as e: if attempt == max_retries - 1: raise e # 最后一次重试失败,抛出异常 time.sleep(retry_delay) retry_delay *= 2 # 指数退避,每次重试间隔翻倍
3. 更换HTTP适配器优化连接
默认的HTTP适配器可能在处理部分连接时存在问题,可以尝试调整连接池参数:
import requests from requests.adapters import HTTPAdapter session = requests.Session() # 增大连接池大小,避免连接耗尽 adapter = HTTPAdapter(pool_connections=10, pool_maxsize=10) session.mount('https://', adapter) response = session.get(url, timeout=10)
4. 排查DNS解析问题
偶发无响应可能和DNS解析有关,可以尝试切换DNS服务器(比如用8.8.8.8),或者测试不同网络环境,排除本地网络的偶发异常。
内容的提问来源于stack exchange,提问作者lhmarsden
相关产品推荐
相关产品推荐

