AWS新手咨询:通过EC2运行Python脚本自动将API数据存入S3是否可行?
Hey there! As someone who's walked plenty of AWS newbies through basic data pipelines, let me start with a clear yes: your plan to run a Python script on EC2 to pull API data and store it in S3 is totally workable. Let's break down the details, including how to stay within the free tier and some small tweaks to make this smoother.
1. 方案可行性分析
Running a Python script on EC2 is a straightforward way to automate this workflow—especially if you're already comfortable writing Python for data fetching. Here's the core flow:
- Use libraries like
requeststo call your third-party API and grab the 3-5MB JSON payload. - Use AWS's official
boto3SDK to upload the JSON directly to your S3 bucket. - Set up a schedule (like Linux
cron) to run the script automatically at intervals you choose (daily, hourly, etc.).
Just remember these key best practices:
- Attach an IAM role to your EC2 instance with permission to write to your S3 bucket. Never hardcode AWS credentials in your script—this is a critical security rule.
- Configure your EC2 security group to allow outbound internet access so the script can reach the third-party API.
2. 如何在AWS免费套餐内操作
AWS's free tier has you covered here, as long as you stick to eligible resources:
- EC2: You get 750 hours per month of a
t2.microort3.microinstance (region-dependent) for free for your first 12 months. That's exactly enough to run one instance 24/7 for a month (31 days = 744 hours, so you'll even have a few hours leftover). - S3: The free tier includes 5GB of standard storage, 20,000 Get Requests, and 2,000 Put Requests per month. With your 3-5MB files, even daily fetches would use a tiny fraction of the 5GB limit.
- Data Transfer: Inbound data to AWS is always free. Since you're only pulling data into EC2 (inbound) and pushing it to S3 (internal AWS network, no outbound charges), you won't hit any transfer fees here.
可选优化:用Lambda替代EC2(更省心,同样免费)
If you want to skip managing an EC2 instance entirely, AWS Lambda is a great serverless alternative. You just upload your Python code, set a schedule via EventBridge, and AWS handles running it. The free tier includes 1 million requests per month and 400,000 GB-seconds of compute time—way more than enough for your small data fetches.
示例Python脚本
Here's a simplified version of your script (install dependencies first on EC2 with pip install requests boto3):
import requests import boto3 import datetime import json def fetch_and_upload_to_s3(): # Fetch data from third-party API api_url = "https://your-third-party-api-endpoint.com/data" response = requests.get(api_url) response.raise_for_status() # Trigger error if API call fails data = response.json() # Initialize S3 client s3 = boto3.client('s3') bucket_name = "your-s3-bucket-name" # Add timestamp to filename to avoid overwriting old data file_key = f"api_data/{datetime.datetime.now().strftime('%Y%m%d_%H%M%S')}.json" # Upload to S3 s3.put_object( Bucket=bucket_name, Key=file_key, Body=json.dumps(data), ContentType='application/json' ) print(f"Successfully uploaded data to S3: {file_key}") if __name__ == "__main__": fetch_and_upload_to_s3()
最后小提醒
- If you're testing, stop your EC2 instance when you're done—even though the free tier covers 750 hours, it's good practice to avoid accidental overuse.
- Enable versioning on your S3 bucket if you want to keep historical copies of your data (this doesn't cost extra in the free tier as long as you stay under 5GB).
内容的提问来源于stack exchange,提问作者Chappleton

