You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python从GitHub读取CSV文件时遇请求耗时过长问题求助

从GitHub Raw链接获取CSV时频繁出现长时间无响应的问题

我需要实时获取GitHub上某CSV文件的最新版本,对应的Raw链接为https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv。文件本身不大,正常1秒内就能完成获取,但尝试了两种导入到Pandas的方法后,都频繁出现长时间卡住无响应的情况(并非每次都会触发)。

尝试过的方法

方法一:直接用Pandas的read_csv

import pandas as pd

url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv'
df = pd.read_csv(url)

经常无报错但长时间运行,最终只能手动终止脚本。

方法二:先用requests获取再导入

import pandas as pd
import requests
import io

url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv'

response = requests.get(url)
df = pd.read_csv(io.StringIO(response.content.decode('utf-8')))

卡在response = requests.get(url)这一行,耗时极长。

已知信息

  • 网络连接正常
  • 不能拉取仓库或下载到本地,必须实时获取最新版本
  • Pandas版本:1.5.2
  • Python版本:3.8.10

解决办法

1. 给请求添加超时时间

不管是requests还是pd.read_csv,都可以设置超时参数,避免无限等待:

  • 针对requests:
response = requests.get(url, timeout=10)  # 设置10秒超时,超时会抛出异常
  • 针对Pandas的read_csv(1.5+版本支持通过storage_options传递参数):
df = pd.read_csv(url, storage_options={'timeout': 10})

2. 增加重试机制

偶发无响应大概率是网络波动或GitHub CDN临时延迟,添加重试逻辑能解决大部分这类问题:

用tenacity库实现重试(需先安装pip install tenacity)

import pandas as pd
import requests
import io
from tenacity import retry, stop_after_attempt, wait_exponential

@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def get_csv_data(url):
    response = requests.get(url, timeout=10)
    response.raise_for_status()  # 触发HTTP错误,比如404、500
    return response.content.decode('utf-8')

url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv'
csv_content = get_csv_data(url)
df = pd.read_csv(io.StringIO(csv_content))

自定义简单重试逻辑(无需额外安装库)

import pandas as pd
import requests
import io
import time

url = 'https://raw.githubusercontent.com/tdwg/rs.tdwg.org/master/terms/terms.csv'
max_retries = 3
retry_delay = 2

for attempt in range(max_retries):
    try:
        response = requests.get(url, timeout=10)
        response.raise_for_status()
        df = pd.read_csv(io.StringIO(response.content.decode('utf-8')))
        break
    except (requests.exceptions.RequestException, pd.errors.ParserError) as e:
        if attempt == max_retries - 1:
            raise e  # 最后一次重试失败,抛出异常
        time.sleep(retry_delay)
        retry_delay *= 2  # 指数退避,每次重试间隔翻倍

3. 更换HTTP适配器优化连接

默认的HTTP适配器可能在处理部分连接时存在问题,可以尝试调整连接池参数:

import requests
from requests.adapters import HTTPAdapter

session = requests.Session()
# 增大连接池大小,避免连接耗尽
adapter = HTTPAdapter(pool_connections=10, pool_maxsize=10)
session.mount('https://', adapter)

response = session.get(url, timeout=10)

4. 排查DNS解析问题

偶发无响应可能和DNS解析有关,可以尝试切换DNS服务器(比如用8.8.8.8),或者测试不同网络环境,排除本地网络的偶发异常。

内容的提问来源于stack exchange,提问作者lhmarsden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 15:02:42