You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas.read_csv读取Google Drive中1000万行大型CSV文件

从Google Drive读取大CSV文件到Pandas DataFrame的解决方案

Hey there! Let's tackle reading that 170MB, 10-million-row CSV from Google Drive into Pandas smoothly. The direct link approach you're using might hit snags with large files or permission checks, so here are a few reliable solutions to get your DataFrame loaded:

1. 修复直接下载链接

Google Drive的默认导出链接对于大文件通常需要额外的确认参数。给你的URL添加&confirm=t,即可绕过下载验证步骤:

import pandas as pd

# 新增确认参数后的链接
follow_network_df = pd.read_csv("https://drive.google.com/uc?export=download&confirm=t&id=1WqH...")

重要提醒:首先要确保你的Google Drive文件已设置为*"任何有链接的人都可以查看"*,否则会遇到权限错误。

2. 使用gdown处理大文件下载

对于像你这种170MB的数据集,gdown库比直接用read_csv更能处理Drive的特殊限制。操作步骤如下:

步骤1:安装gdown

在Jupyter Notebook的终端或代码单元格中运行:

!pip install gdown

步骤2:下载并读取文件

选项1:先保存到本地(适合后续重复使用文件):

import gdown
import pandas as pd

file_id = "1WqH..."
url = f"https://drive.google.com/uc?id={file_id}"
output_file = "follow_network.csv"

# 下载文件
gdown.download(url, output_file, quiet=False)

# 加载为DataFrame
follow_network_df = pd.read_csv(output_file)

选项2:直接读取到内存(不保存本地文件):

import gdown
import pandas as pd
from io import StringIO

file_id = "1WqH..."
url = f"https://drive.google.com/uc?id={file_id}"

# 直接获取文件内容
file_content = gdown.download(url, quiet=False, output=None)

# 转换为可读格式并加载到Pandas
follow_network_df = pd.read_csv(StringIO(file_content.decode('utf-8')))

3. 针对大数据集优化读取参数

因为你处理的是1000万行的数据,可以调整read_csv的参数来节省内存、加快加载速度:

follow_network_df = pd.read_csv(
    "https://drive.google.com/uc?export=download&confirm=t&id=1WqH...",
    dtype={"列名1": str, "列名2": int},  # 指定数据类型,减少内存占用
    low_memory=False,  # 避免混合类型列的警告
    # chunksize=100000  # 如果内存不足,取消注释以分块处理数据
)

如果遇到内存不足的问题,chunksize参数可以让你分批处理数据集,而不是一次性加载全部内容。

内容的提问来源于stack exchange,提问作者Psyduck

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:32:53