如何在AWS Data Pipeline中将EmrActivity转换为HadoopActivity?
Great question! Let's break this down step by step to get your HadoopActivity working exactly like your original EMRActivities.
First: Yes, You Need to Replace the Jar
The DynamoDB EMR Storage Handler jar (s3://dynamodb-emr-<region>/emr-ddb-storage-handler/2.1.0/emr-ddb-2.1.0.jar) you’re currently using is built specifically for reading/writing DynamoDB from Hadoop jobs—it can’t execute spark-submit commands like the command-runner.jar you relied on for your EMRActivities.
You’ll want to stick with command-runner.jar here. This is a built-in EMR utility jar that acts as a wrapper for running cluster commands (including spark-submit), and you can reference it directly by name (no full S3 path required, since EMR clusters have it pre-installed).
Configuring HadoopActivity to Match EMRActivity Behavior
The core difference between EMRActivity and HadoopActivity is how command parameters are structured:
- EMRActivity passes the entire command sequence as a list starting with
command-runner.jar - HadoopActivity splits the jar reference and command arguments into separate fields.
Here’s how to map your original EMRActivity setup to HadoopActivity:
Example Configuration for myHadoopActivity1 (Matching myEmrActivity1)
{ "id": "myHadoopActivity1", "type": "HadoopActivity", "jar": "command-runner.jar", "arguments": [ "spark-submit", "--master", "yarn-cluster", "--deploy-mode", "cluster", "PYTHON=python36", "s3://amznhadoopactivity/school-attendance-python36/calculate_attendance_for_year.py", "#{myYearParameter}" ], "runsOn": {"ref": "myEmrCluster"}, // Link to your existing EMR cluster resource "schedule": {"ref": "mySchedule"} // Link to your pipeline's schedule }
Key Details to Note:
jarfield: Set tocommand-runner.jar(this replaces the first argument in your original EMRActivity)argumentsarray: Include all remaining parameters from your EMRActivity, in the exact same order. This covers:- The
spark-submitcommand itself - Spark configuration flags (
--master,--deploy-mode) - The Python version override (
PYTHON=python36) - Your Spark script’s S3 path
- The parameterized year value (
#{myYearParameter}), just like in your original setup
- The
For myEmrActivity2
Repeat the same pattern: create a second HadoopActivity with jar: "command-runner.jar" and the unique arguments specific to that activity (different script, parameters, etc.).
Quick Additional Checks
- Make sure your EMR cluster has the same Spark and Python 3.6 environment as your original EMRActivities—HadoopActivity runs on the same cluster resources, so environment consistency is critical.
- Pipeline parameters work identically in HadoopActivity, so you can reuse your existing parameter setup without changes.
内容的提问来源于stack exchange,提问作者n1c9

