About us
We have built both — which is why we know what we are replacing
One of us introduced a cloud data warehouse on Snowflake. The other builds AI systems for customers who will not let their data leave the building. lavalake is what follows from both.
Why we are building this
You need a data platform that brings modern tooling with it. The good ones come almost exclusively as a service — and with them you inherit two problems that have nothing to do with each other: a bill that grows with your success, and data that leaves your building.
The bill hits you first. Under credit billing you pay for curiosity: every extra analysis, every dashboard a department opens often, every test of an idea shows up on the next invoice. Sooner or later you throttle data work to throttle cost — the opposite of what you bought the platform for.
Control is about more than location. Who owns the table format when you want to move? Who sees which row when an AI agent asks? Can anything quietly disappear from the log? Anyone accountable for patient records, case files, design or production data decides on a sign-off by those questions — not by the feature list.
And you will have to answer the AI question, probably sooner than planned. It is not whether an agent can query your tables but whether it may, and who can reconstruct afterwards what it saw. That is why the permission chain came first here and the agent second: it gets no account of its own but the rights of whoever is signed in, and every statement is classified, audited and put up for confirmation before it changes anything. An assistant retrofitted later as a feature cannot be given that chain.
None of this is a law of nature. Separated storage and compute, an open table format, a semantic layer and an MCP server all run in your data center. What was missing was a product that ships those layers together — with one permission model instead of six, at a price that does not depend on how curious your analysts are.
We do not come to this from the outside. Data warehouse architecture at IBM, running a cloud data warehouse on Snowflake, twelve years of enterprise systems and AI projects at mindsquare: we know the platforms you want to replace, and the organizations where they are not allowed to run.
Four decisions that explain the product
We did not build a query engine
It runs Trino. Building our own would have been the most expensive and least useful thing to do ourselves — what matters is not who writes the scan but that every statement runs under the exchanged token of whoever is signed in. Same reasoning for Iceberg and Keycloak: open components that do it better, and our work on the chain between them.
A document becomes a table
PDFs do not land in a second content store with its own access control but as chunks in an Iceberg table. The detour sounds awkward and is the reason permissions, audit and lineage apply to documents without a single extra rule.
The search index holds no content
The vector index remembers where something is, not what it says. Hits are re-fetched from Trino under the searcher's token. An index with content would be faster — and a second body of data whose permissions eventually diverge from the real ones.
The BI feed is off as shipped
The outbound data feed answers 403 until an administrator deliberately switches it on. That costs one click and answers the question a sign-off actually asks: which data may leave this building at all.
Where this comes from
Two paths that hit the same wall in 2024 — one from the warehouse side, one from the AI side.
Automotive, from 2016
Data warehouse work at IBM in Analytics, Data & AI. The experience of what a warehouse looks like when it carries production data.
ML in operations, from 2016
On the other side, machine learning inside live operations: ticket classification, predicting incident clusters, automating recurring work — years before generative models drew attention.
Responsibility, from 2019
Fifteen people, twenty-five production applications, ITIL operations with committed service levels. That is where you learn what separates “runs in production” from “runs in the lab”. From 2021, alongside it, research at FernUniversität in Hagen on autonomous multi-agent systems.
Snowflake, 2024
Running a cloud data warehouse on Snowflake per Data Vault 2.0 at SUND Group. After that you know both sides: what such a platform can do, and what it costs.
Self-hosted, 2024
AI agents for customers who will not put their data in a SaaS platform — two platforms of his own, for voice, chat and avatar bots and for autonomous agents, running on self-hosted infrastructure. The same wall, approached from the AI side: the good tools exist only as a service.
The cost question, 2025
The same pattern keeps surfacing in client projects: the bill for the cloud data warehouse grows faster than the value, and nobody can say in advance what an analysis will cost. That names the gap — it is not only a question of data sovereignty but one of predictability.
lavalake, 2026
Lakehouse, governance, search and agent as one product, with one permission model instead of six — on the customer's own hardware.
Who is behind it
lavalake is built by Nils Gregersen and Daniel Alisch. Both write on the blog about the things they work on.

Daniel Alisch
Co-Founder lavalake
Daniel is a Senior AI Consultant at mindsquare AG, where he has come up through every level since 2014: from junior consultant through application ownership to managing consultant, with functional and line responsibility for fifteen people and twenty-five production applications.
He was applying machine learning in operations long before generative models drew everyone's attention: ticket classification, predicting incident clusters, automating recurring operational work. That order — operations first, then the model — shapes lavalake's AI layer.
Today he builds multi-agent systems and RAG pipelines for enterprise environments with LangGraph, MCP, Qdrant and PostgreSQL. Two platforms of his own run in production: one for voice, chat and avatar bots on GDPR-compliant self-hosted infrastructure, and one for autonomous agents in isolated containers with per-task cost attribution — the latter open source.
Twelve years in SAP and non-SAP landscapes brought the part that actually holds data projects up in practice: interfaces. That is why lavalake can attach sources through registered MCP servers — SAP through SAP-MCP — and not only whatever ships a JDBC driver.
The reason was the same one lavalake exists for: one customer would not put data in a SaaS platform, another needed a self-hosted solution. Alongside that, Daniel has been researching autonomous multi-agent systems at FernUniversität in Hagen since 2021.
- Information systems · B.Sc. HWR Berlin, M.Sc. FernUniversität in Hagen
- mindsquare AG · Application owner, ITIL operations with ML
- mindsquare AG · Managing Consultant, 15 people, 25 applications
- SAP · twelve years in SAP and non-SAP landscapes, Basis and interfaces
- FernUniversität in Hagen · research on multi-agent systems, since 2021
- mindsquare AG · Senior AI Consultant, since 2024
More articles by

Nils Gregersen
Co-Founder lavalake
Nils started at Hamburger Hochbahn — an apprenticeship, then deputy and finally full leadership of a team of more than thirty people. He took his information systems degrees part-time at Wismar University of Applied Sciences alongside that, bachelor and master, from 2012 to 2019 — so in parallel with the team lead role and later with his first years at IBM.
What matters for lavalake are the five years after that at IBM, in Analytics, Data & AI: first at the Watson IoT Center on cognitive and analytics projects in the automotive industry, then on the data platform used to train autonomous driving models. That is where you learn what a data platform has to withstand once production and vehicle data come together.
At SUND, as IT Project & AI Innovation Manager since 2024, his remit includes a cloud data warehouse on Snowflake modelled on Data Vault 2.0. So he knows the strengths of such a platform from running one — and the bill, the credit metering and the tie to a closed table format from the same experience.
Then there is the entrepreneurial side: Paycer as co-founder and CEO, then a stint in Singapore as COO of the gold token CACHE and at Silver Bullion, and Spargold since 2026 — environments where traceability is not optional.
- Hamburger Hochbahn · from apprenticeship to team lead, 30+ people, 2009–2016
- Information systems · B.Sc. and M.Sc., Wismar University of Applied Sciences, part-time, 2012–2019
- IBM · five years in Analytics, Data & AI, from 2016
- Paycer · co-founder and CEO, 2021–2023
- CACHE · COO, stint in Singapore
- SUND · IT Project & AI Innovation Manager, since 2024
- Spargold · co-founder and managing director, since 2026
More articles by
- MCP in the data warehouse: agents need a semantic layer, not SQL
- Apache Iceberg on your own hardware: what the table format actually solves
- Separating storage and compute without the cloud: what it takes
- Warehouse migration: what actually takes the time
- Text-to-SQL does not fail on the model, it fails on the metrics
Let us talk about your case
In thirty minutes you will see your architecture, a cost calculation with disclosed assumptions, and an honest assessment of whether the move pays off for you.
Request a demo