← All Updates
Tool LaunchAugust 11, 2026databricks.com 12 reads

Databricks FILE Type Brings Multimodal Data Into Governed Tables

Databricks has introduced a beta FILE column type designed to make documents, images, audio and video easier to query, secure and use in AI workflows.

Databricks has introduced a beta FILE column type designed to make documents, images, audio and video easier to query, secure and use in AI workflows.

Databricks is adding a new native FILE type to its data platform, according to a company blog post, aiming to bring unstructured content into the same managed environment as conventional enterprise tables.

The beta feature lets teams store files such as contracts, product images, call recordings and videos as governed columns, rather than keeping only links to content held elsewhere.

Databricks says the goal is to make multimodal data more usable for AI systems while preserving access controls, compliance processes and portability across the broader data ecosystem.

What Databricks announced

FILE is a new column type for unstructured data. A column type defines what kind of data a table field can hold, such as text, numbers, dates or, in this case, file references.

Instead of putting large binary files directly inside tables, FILE stores lightweight pointers to files in object storage. Binary data means the raw bytes that make up an image, video, document or audio file.

Databricks says this design keeps queries efficient because the platform only pulls the full file content when a workflow explicitly needs it. That matters for video and audio pipelines, where moving large files through every query can slow down processing.

The company also said FILE supports SQL and Python user-defined functions. A user-defined function, or UDF, is custom code that can be run inside a data workflow to transform or analyze information.

Governance for raw files

The central pitch is governance. Databricks says FILE columns can use the same fine-grained access controls and security policies as standard table data, including protections managed through Unity Catalog.

Fine-grained access control means permissions can be applied at a more precise level, such as specific rows or columns, instead of only broad folders or databases. Databricks says FILE supports row-level, column-level and attribute-based access control.

This addresses a common pattern in enterprise AI projects: a table contains a URL or file path, while the actual image, recording or document is controlled by a separate storage system. Databricks argues that this creates two permission models for one dataset.

The company also points to compliance benefits. When a row containing a FILE is deleted, Databricks says the file binary can also be removed from object storage, helping teams handle data-deletion obligations such as GDPR right-to-be-forgotten requests.

Why it matters for multimodal AI

Multimodal AI refers to systems that work across multiple data formats, such as text, images, sound and video. For enterprises, that means AI applications can draw on contracts, inspection images, customer calls and footage, not just structured records.

Databricks gives an example involving an autonomous-driving company investigating why vehicles make unexplained stops. In the scenario, dashcam clips are stored in a FILE column next to trip metadata such as speed and timestamp.

A processing step samples frames from video clips into another FILE column, then an object-detection model adds a hazard label. The resulting table can identify cases where a vehicle stopped despite no visible hazard ahead.

Databricks frames this as a way for agents and analysts to work from a single governed row containing the original file, extracted insights, embeddings and business metadata. An embedding is a numerical representation that helps AI systems compare and retrieve related content.

Open formats and portability

Databricks says it is working with the community to build FILE support into formats and engines including Parquet, Delta Lake, Iceberg and Spark. Parquet, Delta Lake and Iceberg are widely used ways to store and manage analytical data, while Spark is a large-scale data processing engine.

The company positions that work as a safeguard against lock-in, because multimodal data could remain usable across different tools rather than being tied to one vendor or model provider.

Databricks did not announce pricing details for FILE in the source material. The feature is currently described as being in beta, which typically means it is available for testing but may still change before general availability.

What it means for AI learners and prompt users

For PromptsMaze readers, the launch is a reminder that better AI prompts often depend on better data foundations. A prompt asking an agent to analyze video, summarize calls or compare images is only useful if the underlying files are accessible, governed and connected to context.

FILE could make it easier for enterprise teams to build retrieval workflows where prompts draw on documents, images, audio and video together. Retrieval means finding relevant source material for an AI system before it generates an answer.

For learners experimenting with AI agents, the broader lesson is practical: multimodal prompting is not just about model capability. It also depends on how raw evidence is stored, secured, queried and linked to the metadata that gives it meaning.

Source
https://www.databricks.com/blog/introducing-file-type-native-column-type-multimodal-data

Keep exploring

All updates →