TL;DR: Realm now delivers parsed, normalized, and enriched security data directly into Databricks Delta Lake through Zerobus Ingest. The data lands analytics-ready and governed by Unity Catalog, so your team spends its time on analysis, not on preparing data. Because Realm is the managed pipeline doing the work, there is nothing for your team to build or babysit. Realm is among the earliest security partners to natively push structured, enriched security data into Databricks via Zerobus Ingest.
The problem: getting security data into a lakehouse is the hard part
More security teams are moving security telemetry into Databricks. It is where they can hunt across years of history, run ML on their data, and retain everything for a fraction of what a SIEM charges to hold the same volume.
The value is clear. The path in is not. Getting security logs into a lakehouse usually means brittle forwarders, custom parsing jobs, and schema wrangling that breaks every time a vendor changes a log format. Worse, most of that data arrives raw and unstructured, so someone has to build and maintain the jobs that parse and clean it once it lands, on top of the work of moving it in the first place.
The result is a setup that is complex to stand up, slow to change, and fragile to maintain.
How the integration works
Realm sits upstream of Databricks and does the hard part before the data ever hits your lakehouse.
Any source
Cloud, on-prem, or third-party SaaS, using fully managed vendor integrations and the Realm generic collector.
Parse & normalize
Realm parses each log into clean structured JSON and normalizes OCSF observables across every product.
In real time
Enriched with GeoIP, threat intel, and your own context from CMDB or HRIS, as the event flows through.
Into Databricks
Realm delivers both the raw logs and the parsed, normalized version, so you decide what lands where across your medallion architecture.
Realm collects logs from any source, cloud, on-prem, or third-party SaaS, using fully managed vendor integrations and the Realm generic collector. As data flows through, Realm parses it into clean structured JSON, extracts and normalizes OCSF observables across every product, and enriches it in real time with datasets like GeoIP, threat intel, and your own business context from CMDB or HRIS.
From there, Realm streams the data directly into the Unity Catalog Delta tables you choose, through Zerobus Ingest, Databricks’ serverless push-based ingestion API. Realm delivers both the raw, unparsed logs and the parsed, enriched, and normalized version, so you decide where each lands: raw logs into your Bronze layer, structured data into Silver, or both, whatever your medallion architecture calls for. Because the parsed data arrives already clean and normalized, your downstream Silver and Gold transformations are simpler and faster to build.
Direct write, not stage-and-load
Most tools that write to Databricks stage data to a cloud storage bucket first, then load it into the lakehouse in batches. That means an extra storage tier to manage, added latency, and more moving parts to break.
Realm skips it. Zerobus Ingest is a push-based API that writes directly into Unity Catalog Delta tables, with no customer-managed staging bucket or message bus. Realm streams structured events straight to the table in near real time, and Zerobus Ingest scales simply by opening more connections. Fewer parts, lower latency, less to maintain. And because Zerobus Ingest writes into Unity Catalog, every record inherits Databricks governance, lineage, and access controls the moment it lands.
Why it matters
No more plumbing. Realm’s managed integrations and collector replace the fragile forwarders and custom parsing jobs teams assemble to feed Databricks. You configure the output feed in the Realm UI and data starts flowing. That is the whole setup.
Less engineering effort. Because Realm forwards parsed, structured, enriched data instead of raw logs, your team skips the work of building and maintaining parsing and cleanup jobs downstream. The transformation happens once, in Realm, before the data lands, so your engineers spend their time on detection and analysis, not data prep.
Actionable on arrival. OCSF-normalized observables and point-in-time enrichment mean the data is ready to query, correlate, and model the moment it hits the lakehouse. Enrichment happens in real time as data moves through Realm, so you capture the values as they were at the moment of the event, not whatever they resolve to days later. After a one-time setup in Databricks (creating the target table, Zerobus Ingest endpoint, and access credentials), the data is queryable the moment it lands, with no additional transformation needed before you can use it.
What this looks like for a Realm + Databricks customer
One global enterprise customer is adopting the integration to replace exactly the kind of setup described above: a complex, fragile, and clunky pipeline they built themselves to get security logs into Databricks.
With Realm, that self-built plumbing goes away entirely. And because Realm forwards parsed, structured data rather than raw logs, they expect to save their team significant time and effort, on top of the savings Realm already delivers on ingest. Just as important, cleaner data in the lakehouse opens up new use cases they could not easily support before.
Where this is going
High-throughput, direct-to-lakehouse ingestion is just the foundation. Looking ahead, this Zerobus Ingest integration will directly accelerate data collection into Lakewatch (currently in private preview), drastically reducing the complexity and time required to feed high-fidelity security telemetry into detection and response workflows.
For security teams looking to modernize their SIEM strategy on Databricks, the combination of Realm and Lakewatch eliminates the traditional log transport and parsing. Teams can easily migrate away from costly, legacy SIEM architectures toward a lean, unified data intelligence platform on Databricks without compromising on speed or coverage.
Get started
If Databricks is where your security data belongs, Realm gets it there clean, enriched, and ready to use, without building and maintaining the pipeline yourself.
Read the integration docs for step-by-step setup, or talk to our team to see it on your own data.