MCP & AI
From question to dashboard
Query metrics and create dashboards in natural language. Your own agents can connect through MCP too.
- Choose your own model
- Changes require your approval
Query your business data with AI, SQL and dashboards. Keep data, search and access permissions on your own infrastructure.
AI in practice
The assistant can explore data, query defined metrics and create dashboards. You work in natural language; tools execute each step under your identity. This example illustrates a possible workflow using a previously defined revenue metric.
Read the docsHow is our revenue developing by region?
The assistant queries a previously defined metric.
Create a dashboard from this.
Review the proposed change and confirm it.
One view for your team
Each tile runs under the permissions of the person viewing it.
The AI data platform
Connect data, find answers and control access in one platform.
MCP & AI
Query metrics and create dashboards in natural language. Your own agents can connect through MCP too.
Governance
People and agents work under the same permissions. Access and changes remain traceable.
Getting data in
Bring databases, files and documents into one lakehouse. Guided imports help you get started.
Knowledge & metrics
Use the same defined metrics across search, SQL, dashboards and AI answers.
Architecture
One web console for data imports, SQL, search, AI assistance and dashboards. Underneath sits an open Iceberg lakehouse with separate storage and compute. Run the platform in your data center, choose your AI model and keep your data in open formats.
The reference deployment runs as a Docker Compose stack on one machine; Ansible prepares installation and updates on Debian/Ubuntu. The Helm path for Kubernetes still has operational gaps, including policy integration. High availability and fully air-gapped operation remain planned.
Console (SPA) · SQL worksheet · dashboards · AI assistant · MCP server — all behind one reverse proxy
Keycloak OIDC · token exchange per RFC 8693 · row-access and masking policies · WORM audit · lineage
Trino under your identity · S, M and L workload classes in the Compose stack
Apache Iceberg with Nessie: ACID · time travel · branching · schema evolution · open to Spark, Flink, Trino
S3-compatible (MinIO, Ceph, NetApp StorageGRID, Dell ECS) · PostgreSQL with pgvector as the vector index
The concepts stay, the bill changes. What is still missing says so.
Cost
Predictable platform costs instead of credits per query. See what you could save with lavalake on your own infrastructure, including hardware, power and operations.
No credits
A fixed subscription based on vCPU and support instead of query credits. Queries from people and agents use your compute capacity without metering each operation.
No egress
lavalake charges no egress fees on your data. Budget separately for your infrastructure, network connectivity and connected providers.
Hardware you already own
lavalake runs on your existing Kubernetes and your S3-compatible storage. Capacity you have already written off keeps earning.
Planning you can budget
Cost follows cluster size and term, not usage behavior. Your teams are free to explore the data again.
Model assumption: 6 servers, 48 vCPU · Business
Review assumptions and costs →Savings over 3 years
€1,304,343
Model of recurring costs. We calculate migration, parallel operation, floor space and AI models separately for your setup.
Request a TCO analysisroughly half
Andreessen Horowitz estimates that repatriating cloud workloads cuts total cost of ownership to about half. An estimate rather than a survey, and calculated for software companies at scale.
a16z, The Cost of Cloud, a Trillion Dollar Paradox, 2021$2 million a year
37signals published the breakdown of its cloud exit: $3.2 million in annual spend, then roughly $2 million less per year after $700K of hardware. A single case with real invoices — and not a data warehouse.
37signals, Cloud Exit, 2024–2025Sovereignty
Giving AI access to business data means controlling permissions and where data goes. In lavalake, people and agents share the same access rules. You choose which model to connect: self-hosted on your network or with an external provider.
Use cases
Assistants find documents, query curated metrics and create dashboards. Your own agents can access these tools through MCP — under the user’s permissions, with actions recorded in the audit log.
Search, metrics and dashboards through MCP
Production telemetry meets SAP orders — without mirroring sensor data into a public cloud. Scrap analysis runs on the same cluster as financial controlling.
Analyze production and ERP data together
Regulatory reporting and risk models read from versioned Iceberg tables. Time travel reproduces any reporting date, and lineage answers the supervisor's question.
Trace historical data snapshots
Treatment and billing data stay in-house, with analysis and AI assistance running on top. Roles and masking make sure each department sees only its own slice.
Masking applies in the engine, not in the interface
Store data comes into the lakehouse from the upstream systems, and replenishment and marketing compute on the same curated metrics of the semantic layer instead of each on its own extract.
One metric definition for every report
Data from separate case management systems is brought together without leaving the legal framework. Operated in the state data center, analyzed by self-service.
Operated in the state data center
Migration
Migration here does not mean rebuilding. Tables come over as Iceberg, roles and permissions move into Keycloak and the Trino policies. Both systems run in parallel until the results reconcile.
We read your existing warehouse metadata: tables, views, roles, query history and cost distribution. The result is a migration plan with a defensible TCO calculation.
lavalake is installed on your cluster via Helm chart and wired to your S3-compatible storage, your AD and your monitoring. No new hardware purchase required.
Tables are taken over as Iceberg and queries are re-pointed at Trino. Both systems run in parallel until the results match row for row.
BI connections point at lavalake, roles and row filters are in place, reports are validated against the old figures. Then the cloud subscription gets canceled.
Pricing
The AI assistant, MCP server, search and governance are included in every tier. Pricing follows allocated vCPU and support, not the number of data queries. Budget separately for external AI providers or self-hosted models.
Your data lives on your storage. The subscription depends on allocated vCPU and support; data volume is not billed separately.
€249
per month, up to 8 vCPU
For a department's first production warehouse: the complete feature set on a single machine, and someone who answers.
€549
per month, up to 24 vCPU
When the warehouse outgrows the department: three times the compute, and no user limit.
€1,250
per month, up to 64 vCPU
For several departments on one platform: the cluster grows, the user count is open, the response time is committed.
€2,499
per month, per cluster
For production use across several departments: unlimited compute, and tickets that are handled before everyone else's.
Custom
Regulated environments
For organizations under special obligations: separated operation, offline updates, source code escrow and on-site support.
60 million rows and around 50 simulated analysts on a machine with 6 vCPU and 24 GB RAM. Our documented load test shows what compact hardware can do. The measurements in detail
All prices net, excluding VAT. What counts are the vCPUs allocated to the platform's services, not your servers as a whole. There are no hosting costs — the platform runs on your infrastructure.
DocumentationFrom the first start on your own machine through the concepts to the switches that matter in production.
15 pages
BlogWhy cloud warehouses get expensive, what sovereignty means technically, and why agents need a semantic layer.
10 articles
CalculatorA model from disclosed assumptions — each with a source or explicitly labelled an estimate, all adjustable.
9 assumptions, sourced one by one
In 30 minutes, we show how the assistant finds data, queries metrics and creates dashboards. Together, we discuss deployment in your data center, your choice of model and the right access permissions.