Skip to main content

ETL pipeline with Amazon Redshift and AWS Glue

An ETL pipeline with Amazon Redshift and AWS Glue

This example lives in the pulumi/examples repository. Check out just this directory to use it:

Get started with this example
git clone --filter=blob:none --sparse https://github.com/pulumi/examples pulumi-examples
git -C pulumi-examples sparse-checkout set aws-py-redshift-glue-etl
cd pulumi-examples/aws-py-redshift-glue-etl

This example creates an ETL pipeline using Amazon Redshift and AWS Glue. The pipeline extracts data from an S3 bucket with a Glue crawler, transforms it with a Python script wrapped in a Glue job, and loads it into a Redshift database deployed in a VPC.

Prerequisites#

  1. Install Pulumi.
  2. Install Python.
  3. Configure your AWS credentials.

Deploying the App#

  1. Clone this repo, change to this directory, then create a new stack for the project:

    Terminal window
    pulumi stack init
  2. Specify an AWS region to deploy into:

    Terminal window
    pulumi config set aws:region us-west-2
  3. Install Python dependencies and run Pulumi:

    Terminal window
    python3 -m venv venv
    source venv/bin/activate
    pip install -r requirements.txt
    pulumi up
  4. In a few moments, the Redshift cluster and Glue components will be up and running and the S3 bucket name emitted as a Pulumi stack output.

    Terminal window
    ...
    Outputs:
    dataBucketName: "events-56e424a"
  5. Upload the included sample data file to S3 to verify the automation works as expected:

    Terminal window
    aws s3 cp events-1.txt s3://$(pulumi stack output dataBucketName)
  6. When you’re ready, destroy your stack and remove it:

    Terminal window
    pulumi destroy --yes
    pulumi stack rm --yes

Related

The infrastructure as code platform for any cloud.