You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests.get()获取S3页面仅得预加载页,如何获取真实内容?

问题原因

这个页面是AWS S3托管的静态页面,它的文件列表是通过JavaScript动态渲染的。requests.get()只会获取初始的HTML内容,不会执行页面里的JS代码,所以你拿到的只是负责加载真实内容的框架页面,看不到实际的zip文件链接。

解决办法

提供两种实用方案,选你顺手的来:

方案1:用boto3直接访问S3存储桶(推荐)

S3存储桶本身支持程序式访问,不需要解析网页。步骤如下:

  1. 安装boto3:
pip install boto3
  1. 编写脚本列出所有zip文件并下载:
import boto3
from botocore import UNSIGNED
from botocore.config import Config

# 初始化无需认证的S3客户端(这个桶是公开可读的)
s3 = boto3.client('s3', config=Config(signature_version=UNSIGNED))

# 指定存储桶名称
bucket_name = 'divvy-tripdata'

# 列出桶里所有zip文件
response = s3.list_objects_v2(Bucket=bucket_name)
for obj in response.get('Contents', []):
    key = obj['Key']
    if key.endswith('.zip'):
        print(f"下载文件: {key}")
        # 下载文件到本地,文件名和桶里的一致
        s3.download_file(bucket_name, key, key)

这个方法更可靠,直接和S3的API交互,不会受页面渲染逻辑变化影响。

方案2:用支持JS渲染的工具抓取页面

如果不想用AWS SDK,可以用requests-html这类能执行JS的库:

  1. 安装requests-html:
pip install requests-html
  1. 编写脚本:
from requests_html import HTMLSession
import requests

session = HTMLSession()
url = 'https://divvy-tripdata.s3.amazonaws.com/index.html'

# 获取页面并执行JS
r = session.get(url)
r.html.render()  # 执行页面JS,渲染出真实内容

# 提取所有zip链接
zip_links = r.html.find('a[href$=".zip"]')
for link in zip_links:
    zip_url = link.attrs['href']
    # 处理相对链接,拼接成完整URL
    if not zip_url.startswith('http'):
        zip_url = f"https://divvy-tripdata.s3.amazonaws.com/{zip_url}"
    print(f"下载: {zip_url}")
    # 下载文件
    response = requests.get(zip_url)
    filename = zip_url.split('/')[-1]
    with open(filename, 'wb') as f:
        f.write(response.content)

这个方法模拟浏览器执行JS,能拿到和浏览器里看到的一致的页面内容。

内容的提问来源于stack exchange,提问作者StandardIO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 18:16:08