Databricks VACUUM removes obsolete data files that are no longer referenced by a Delta table and have passed the configured retention period. It helps control cloud storage costs, but it also removes the data files required for older Delta Lake table versions. This guide explains the VACUUM command, retention settings, DRY RUN, safe usage, and how to vacuum multiple tables.
What is VACUUM in Delta Lake?
A Delta table stores data files and a transaction log that tracks table versions. When data is updated or deleted, older data files may no longer be referenced by the current table state. VACUUM permanently removes eligible files from storage after the retention threshold has expired.
The default VACUUM retention threshold is 7 days. The retention period is controlled by the delta.deletedFileRetentionDuration table property. A longer retention period preserves more historical data but uses more storage.
Basic VACUUM syntax
VACUUM catalog.schema.table_name;
For a path-based Delta table, use:
VACUUM delta.`abfss://container@storageaccount.dfs.core.windows.net/path/to/table`;
Preview files with VACUUM DRY RUN
Use DRY RUN to preview files eligible for deletion without deleting them. Databricks returns a list of up to 1,000 files.
VACUUM catalog.schema.table_name DRY RUN;
Review the result before running the actual VACUUM operation, especially on production tables or tables used by long-running jobs.
Configure the retention period
You can increase the retention period by setting the Delta table property. For example, the following command retains eligible deleted files for 30 days:
ALTER TABLE catalog.schema.table_name
SET TBLPROPERTIES (
'delta.deletedFileRetentionDuration' = '30 days'
);
After changing the property, run VACUUM normally:
VACUUM catalog.schema.table_name;
Important warning about short retention periods
Databricks strongly recommends a retention interval of at least 7 days. A short retention period can delete files still needed by long-running or concurrent jobs. It also reduces the time-travel history available for the table.
Do not disable the retention safety check in production unless you have verified that no operation can run longer than the selected retention period. Disabling the check can cause data loss or failed reads.
# Use only for a controlled test after confirming that no long-running
# or concurrent operation depends on files within the retention window.
spark.conf.set("spark.databricks.delta.retentionDurationCheck.enabled", "false")
Never use a zero-hour retention period casually. It can remove files required by active jobs and older table versions. Prefer the default retention or a deliberately selected retention period that matches your recovery and time-travel requirements.
Run VACUUM on multiple Delta tables
The following PySpark example discovers tables in a schema and runs VACUUM for each table. Test it in a non-production environment first, and avoid changing retention settings automatically unless that is part of your approved data-retention policy.
from pyspark.sql import functions as F
catalog_name = "main"
schema_name = "silver"
tables = (
spark.sql(f"SHOW TABLES IN `{catalog_name}`.`{schema_name}`")
.select("tableName")
.collect()
)
for row in tables:
table_name = row["tableName"]
full_table_name = f"`{catalog_name}`.`{schema_name}`.`{table_name}`"
print(f"Running VACUUM on {full_table_name}")
spark.sql(f"VACUUM {full_table_name}")
For a safer review workflow, run DRY RUN first for the specific tables you intend to clean. Also verify that the selected tables are Delta tables and that their retention settings meet your organization’s recovery requirements.
VACUUM and Delta Lake time travel
VACUUM removes data files that are no longer needed by table versions within the retention window. As a result, you may lose the ability to time travel to versions older than the configured data-file retention period.
SELECT *
FROM catalog.schema.table_name
VERSION AS OF 10;
The query above may fail if the data files required for version 10 have already been removed. Delta log retention and data-file retention are related but different: VACUUM removes obsolete data files, while transaction-log cleanup is handled separately.
VACUUM FULL and LITE modes
On supported Databricks Runtime versions, FULL is the default mode. It scans the table directory and removes eligible files, including files not referenced by the transaction log. LITE uses the Delta transaction log to identify files and can reduce directory-listing work, but it cannot remove files that are not represented in the log. If LITE cannot safely complete, use FULL.
VACUUM catalog.schema.table_name FULL;
VACUUM catalog.schema.table_name LITE;
Best practices
- Use the default 7-day retention unless your workload requires a different period.
- Run
DRY RUNbefore cleanup when validating a new table or retention policy. - Do not run VACUUM while long-running jobs may still need files inside the retention window.
- Keep retention aligned with time-travel, recovery, audit, and compliance requirements.
- Use approved scheduling or platform-managed optimization where available.
- Monitor the number of files removed and the duration of large VACUUM operations.
- Do not store unrelated files inside a Delta table directory unless you understand how VACUUM treats them.
Related Microsoft Fabric topic: If you are working with Delta files in a Fabric Lakehouse, see How to Delete Stale Delta Files in Microsoft Fabric Lakehouse and compare its cleanup approach with Databricks VACUUM.
Conclusion
Databricks VACUUM is a useful maintenance command for removing obsolete Delta data files and controlling storage growth. Use it carefully: preview eligible files, select a retention period deliberately, and remember that removing old data files can limit time travel to historical table versions.
See more
Visual Studio Marketplace
SSIS Catalog Migration Wizard
Extend Visual Studio with an easy way to migrate SSIS Catalog projects.
Kunal Rathi
With over 15 years of experience in data engineering and analytics, I've assisted countless clients in gaining valuable insights from their data. As a dedicated supporter of Data, Cloud and DevOps, I'm excited to connect with individuals who share my passion for this field. If my work resonates with you, we can talk and collaborate.






