Spark on Azure HDInsight
Spark on Azure HDInsight example in Python
This example lives in the pulumi/examples repository. Check out just this directory to use it:
git clone --filter=blob:none --sparse https://github.com/pulumi/examples pulumi-examplesgit -C pulumi-examples sparse-checkout set classic-azure-py-hdinsight-sparkcd pulumi-examples/classic-azure-py-hdinsight-sparkAn example Pulumi component that deploys a Spark cluster on Azure HDInsight.
Running the App#
-
Create a new stack:
Terminal window pulumi stack init dev -
Login to Azure CLI (you will be prompted to do this during deployment if you forget this step):
Terminal window az login -
Specify the Azure location and subscription to use:
Terminal window pulumi config set azure:location WestUSpulumi config set azure:subscriptionId <YOUR_SUBSCRIPTION_ID> -
Define Spark username and password (make it complex enough to satisfy Azure policy):
Terminal window pulumi config set username <value>pulumi config set --secret password <value> -
Run
pulumi upto preview and deploy changes:Terminal window pulumi upPreviewing changes:...Performing changes:...info: 5 changes performed:+ 5 resources createdUpdate duration: 15m6s -
Check the deployed Spark endpoint:
Terminal window pulumi stack output endpointhttps://myspark1234abcd.azurehdinsight.net/# For instance, Jupyter notebooks are available at https://myspark1234abcd.azurehdinsight.net/jupyter/# Follow https://docs.microsoft.com/en-us/azure/hdinsight/spark/apache-spark-load-data-run-query to test it out