Quick answer: Terraform can manage Azure Databricks infrastructure and workspace resources as code. Use the AzureRM provider to create or manage the Azure Databricks workspace, and the Databricks Terraform provider to manage resources inside the workspace such as clusters, jobs, users, groups, permissions, and notebooks.

Prerequisites:
1. Terraform installed and available in your PATH.
2. An Azure subscription and permission to create or manage the Databricks workspace.
3. An Azure authentication method that Terraform can use, such as Azure CLI authentication, a service principal, or managed identity.
4. The Databricks Terraform provider configured for the workspace resources you plan to manage.

1. Configure Terraform providers

Start by declaring the AzureRM provider for Azure resources and the Databricks provider for resources inside the Databricks workspace. The configuration below tells Terraform where to download both providers from.

terraform { required_providers { azurerm = { source = “hashicorp/azurerm” } databricks = { source = “databricks/databricks” } } } provider “azurerm” { features {} }

Save the configuration as a Terraform file such as main.tf. Then run terraform init to download the required providers and initialize the working directory.

2. Create an Azure Databricks workspace

The AzureRM provider manages the Azure-side resources. In this example, Terraform first creates a resource group and then creates an Azure Databricks workspace inside it. The sku determines the workspace pricing tier.

resource “azurerm_resource_group” “example” { name = “example-databricks-rg” location = “East US” } resource “azurerm_databricks_workspace” “example” { name = “example-databricks” resource_group_name = azurerm_resource_group.example.name location = azurerm_resource_group.example.location sku = “premium” }

The azurerm_databricks_workspace resource manages the Azure Databricks workspace itself. After Terraform creates the workspace, its workspace_url can be used to configure the Databricks provider.

3. Configure the Databricks provider

The Databricks provider needs to know which workspace it should manage. The AzureRM workspace resource exposes the workspace URL, so it can be passed directly to the Databricks provider.

provider “databricks” { host = azurerm_databricks_workspace.example.workspace_url }

This example only shows the workspace endpoint. For automation, configure an appropriate non-interactive authentication method such as a service principal or managed identity rather than storing credentials directly in the Terraform configuration.

4. Manage Databricks resources

Once the Databricks provider is connected to the workspace, you can manage resources inside the workspace. The example below uses data sources to select an available long-term-support Spark version and a node type with local disk, then creates a small cluster.

data “databricks_spark_version” “latest_lts” { long_term_support = true } data “databricks_node_type” “smallest” { local_disk = true } resource “databricks_cluster” “example” { cluster_name = “example-cluster” spark_version = data.databricks_spark_version.latest_lts.id node_type_id = data.databricks_node_type.smallest.id autotermination_minutes = 20 num_workers = 2 }

Here, the two data blocks query Databricks for values that already exist, while databricks_cluster defines the cluster Terraform should create. Adjust the cluster size, Spark version, and auto-termination setting for your workload rather than copying these values directly into production.

5. Manage jobs and permissions

The Databricks provider can also manage jobs and other workspace resources. For example, a job can reference a cluster and run a notebook on a schedule. The exact task configuration depends on how your Databricks workloads are organized.

resource “databricks_job” “example” { name = “example-job” task { task_key = “example-task” new_cluster { spark_version = data.databricks_spark_version.latest_lts.id node_type_id = data.databricks_node_type.smallest.id num_workers = 1 } notebook_task { notebook_path = “/Shared/example-notebook” } } }

This creates a simple Databricks job with a task that runs a notebook on a new job cluster. The same provider can also be used for users, groups, notebooks, permissions, and other supported Databricks resources.

6. Deploy with CI/CD

Store the Terraform configuration in source control and validate the configuration before applying changes. A typical workflow is:

terraform fmt
terraform validate
terraform plan
terraform apply

terraform fmt formats the configuration, terraform validate checks the configuration syntax and structure, and terraform plan shows the changes Terraform intends to make. Run terraform apply after the changes have been reviewed and approved. For team-based CI/CD, keep the Terraform state in a suitable remote backend rather than relying on a local state file.

Pro tips:
Keep Azure infrastructure resources and Databricks workspace resources conceptually separate. Use the AzureRM provider for the Azure Databricks workspace and the Databricks provider for resources inside the workspace.

See more

Visual Studio Marketplace

SSIS Catalog Migration Wizard

Extend Visual Studio with an easy way to migrate SSIS Catalog projects.

James Sandy

James is interested in solving problems using data, AI, and cloud technology, and he researches AI safety.