Skip to content
MCP server built inHow that works
Cloud convenience. Your infrastructure.

Data cloud for AI & analytics.
In your own data center.

Query your business data with AI, SQL and dashboards. Keep data, search and access permissions on your own infrastructure.

AI
Ask questions, find knowledge, create dashboards
MCP
Connect your own agents to your data

AI in practice

What your assistant can do with lavalake

Ask questions. Put your data to work.

The assistant can explore data, query defined metrics and create dashboards. You work in natural language; tools execute each step under your identity. This example illustrates a possible workflow using a previously defined revenue metric.

Read the docs
  1. 1

    Ask a question

    How is our revenue developing by region?

    The assistant queries a previously defined metric.

  2. 2

    Approve a dashboard

    Create a dashboard from this.

    Review the proposed change and confirm it.

  3. 3

    Analyze together

    One view for your team

    Each tile runs under the permissions of the person viewing it.

The AI data platform

Built for agents, not retrofitted.

Connect data, find answers and control access in one platform.

MCP & AI

From question to dashboard

Query metrics and create dashboards in natural language. Your own agents can connect through MCP too.

  • Choose your own model
  • Changes require your approval
Learn more: From question to dashboard

Getting data in

Make business data available to AI

Bring databases, files and documents into one lakehouse. Guided imports help you get started.

  • Import databases, CSV, PDF and Word
  • Continuous ingestionIn development
Learn more: Make business data available to AI

Architecture

Cloud convenience for data and AI. On your infrastructure.

One web console for data imports, SQL, search, AI assistance and dashboards. Underneath sits an open Iceberg lakehouse with separate storage and compute. Run the platform in your data center, choose your AI model and keep your data in open formats.

The reference deployment runs as a Docker Compose stack on one machine; Ansible prepares installation and updates on Debian/Ubuntu. The Helm path for Kubernetes still has operational gaps, including policy integration. High availability and fully air-gapped operation remain planned.

Explore the architecture and technical details
One platform, one authorization model: every access — human, agent or BI tool — runs through the same identity and the same audit trail.
01

Access

Console (SPA) · SQL worksheet · dashboards · AI assistant · MCP server — all behind one reverse proxy

02

Identity & governance

Keycloak OIDC · token exchange per RFC 8693 · row-access and masking policies · WORM audit · lineage

03

Compute

Trino under your identity · S, M and L workload classes in the Compose stack

04

Catalog & table format

Apache Iceberg with Nessie: ACID · time travel · branching · schema evolution · open to Spark, Flink, Trino

05

Storage & index

S3-compatible (MinIO, Ceph, NetApp StorageGRID, Dell ECS) · PostgreSQL with pgvector as the vector index

If you already know a cloud warehouse

The concepts stay, the bill changes. What is still missing says so.

Proprietary table format
Apache Iceberg with Nessie (open, versioned)
Web interface for queries
Console with SQL worksheet and data browser
Roles and row filters
Policies in Trino, under delegated identity
Access history and provenance
WORM audit log, lineage and access history
Full-text and vector search
Hybrid search with a local embedding model
The vendor's AI features
MCP server · the model of your choice
Virtual warehouse
Elastic compute clustersIn development
Scheduled tasks
Schedules for reports, threshold alerts and Iceberg maintenance
Streams
Continuous ingest and CDCIn development
Credit billing
A fixed subscription

Cost

More questions. More AI. No credits per data query.

Predictable platform costs instead of credits per query. See what you could save with lavalake on your own infrastructure, including hardware, power and operations.

€3,000€300,000

No credits

A fixed subscription based on vCPU and support instead of query credits. Queries from people and agents use your compute capacity without metering each operation.

No egress

lavalake charges no egress fees on your data. Budget separately for your infrastructure, network connectivity and connected providers.

Hardware you already own

lavalake runs on your existing Kubernetes and your S3-compatible storage. Capacity you have already written off keeps earning.

Planning you can budget

Cost follows cluster size and term, not usage behavior. Your teams are free to explore the data again.

Monthly comparison

81% cheaper
Cloud warehouse today€45,000
lavalake on-premises€8,768
lavalake subscription
€1,250
Hardware amortization
€3,500
Operations & administration
€1,913
Power, including cooling and losses
€812
Storage, network, redundancy, backup
€1,294

Model assumption: 6 servers, 48 vCPU · Business

Review assumptions and costs →

Savings over 3 years

€1,304,343

Model of recurring costs. We calculate migration, parallel operation, floor space and AI models separately for your setup.

Request a TCO analysis
Sources for context
  • roughly half

    Andreessen Horowitz estimates that repatriating cloud workloads cuts total cost of ownership to about half. An estimate rather than a survey, and calculated for software companies at scale.

    a16z, The Cost of Cloud, a Trillion Dollar Paradox, 2021
  • $2 million a year

    37signals published the breakdown of its cloud exit: $3.2 million in annual spend, then roughly $2 million less per year after $700K of hardware. A single case with real invoices — and not a data warehouse.

    37signals, Cloud Exit, 2024–2025

Sovereignty

AI on sensitive data. With clear boundaries.

Giving AI access to business data means controlling permissions and where data goes. In lavalake, people and agents share the same access rules. You choose which model to connect: self-hosted on your network or with an external provider.

You choose where your AI runs

  • The lakehouse and local search stay on your network; external AI providers receive conversation content and tool results used by the assistant
  • The vendor has no access to content or metadata
  • The embedding model runs locally; semantic search needs no cloud key
  • Fully air-gapped operation, including installation and updatesIn development

Provable, not asserted

  • WORM audit log as a hash chain, over users and agents alike
  • Lineage from source to reported figure, access history per table
  • Every query under the identity of whoever is signed in — no shared service account
  • Table and column classification with confirmed suggestions, without sending data values to the model
  • Personal-data access, deletion and retention workflowsIn development

Operations

  • Docker Compose for one machine, Ansible for installation and updates; Helm as the Kubernetes path
  • Everything behind one reverse proxy; Trino and storage are unreachable from outside
  • Daily database and configuration backup in the Compose stack; lakehouse files need separate backup
  • High availability, a central observability stack and secrets through VaultIn development
Governance in detail →

Use cases

AI and analytics where your data lives

Request a reference architecture
AI applications

From business knowledge to answers

Assistants find documents, query curated metrics and create dashboards. Your own agents can access these tools through MCP — under the user’s permissions, with actions recorded in the audit log.

Search, metrics and dashboards through MCP

Manufacturing

Machine data and ERP in one analysis

Production telemetry meets SAP orders — without mirroring sensor data into a public cloud. Scrap analysis runs on the same cluster as financial controlling.

Analyze production and ERP data together

Financial services

Risk calculation with auditable provenance

Regulatory reporting and risk models read from versioned Iceberg tables. Time travel reproduces any reporting date, and lineage answers the supervisor's question.

Trace historical data snapshots

Healthcare

Clinical analysis inside your own network

Treatment and billing data stay in-house, with analysis and AI assistance running on top. Roles and masking make sure each department sees only its own slice.

Masking applies in the engine, not in the interface

Retail

One metric instead of four exports

Store data comes into the lakehouse from the upstream systems, and replenishment and marketing compute on the same curated metrics of the semantic layer instead of each on its own extract.

One metric definition for every report

Public sector

Making case management systems analyzable

Data from separate case management systems is brought together without leaving the legal framework. Operated in the state data center, analyzed by self-service.

Operated in the state data center

Migration

Four weeks, and your cloud subscription is last winter's problem

Migration here does not mean rebuilding. Tables come over as Iceberg, roles and permissions move into Keycloak and the Trino policies. Both systems run in parallel until the results reconcile.

  1. Week 1

    Taking stock

    We read your existing warehouse metadata: tables, views, roles, query history and cost distribution. The result is a migration plan with a defensible TCO calculation.

  2. Week 2

    Installation

    lavalake is installed on your cluster via Helm chart and wired to your S3-compatible storage, your AD and your monitoring. No new hardware purchase required.

  3. Week 3

    Data and pipelines

    Tables are taken over as Iceberg and queries are re-pointed at Trino. Both systems run in parallel until the results match row for row.

  4. Week 4

    Cutover

    BI connections point at lavalake, roles and row filters are in place, reports are validated against the old figures. Then the cloud subscription gets canceled.

Pricing

A subscription, not a meter

The AI assistant, MCP server, search and governance are included in every tier. Pricing follows allocated vCPU and support, not the number of data queries. Budget separately for external AI providers or self-hosted models.

Your data lives on your storage. The subscription depends on allocated vCPU and support; data volume is not billed separately.

Team

€249

per month, up to 8 vCPU

For a department's first production warehouse: the complete feature set on a single machine, and someone who answers.

  • Iceberg lakehouse on your storage
  • SQL worksheet, import assistant and dashboards
  • MCP server, hybrid search and metrics
  • Up to 8 vCPU
  • Up to 10 users
  • Ticket support, answer within a business day
Request a trial

Growth

€549

per month, up to 24 vCPU

When the warehouse outgrows the department: three times the compute, and no user limit.

  • The same feature set as Team
  • Up to 24 vCPU
  • Unlimited users
  • Ticket support, answer within a business day
Request a quote

Business

€1,250

per month, up to 64 vCPU

For several departments on one platform: the cluster grows, the user count is open, the response time is committed.

  • The same feature set as Team
  • Up to 64 vCPU
  • Unlimited users
  • 8/5 support with a committed response time
Request a quote
Recommended

Enterprise

€2,499

per month, per cluster

For production use across several departments: unlimited compute, and tickets that are handled before everyone else's.

  • The same feature set as Team
  • Unlimited vCPU
  • 8/5 support with a committed response time
  • Prioritized ticket handling
Request a quote

Sovereign

Custom

Regulated environments

For organizations under special obligations: separated operation, offline updates, source code escrow and on-site support.

  • Operation without an internet connectionIn development
  • Secrets through Vault or External SecretsIn development
  • Source code escrow
  • On-site engagements
  • Help with your certifications
Schedule a call

60 million rows and around 50 simulated analysts on a machine with 6 vCPU and 24 GB RAM. Our documented load test shows what compact hardware can do. The measurements in detail

All prices net, excluding VAT. What counts are the vCPUs allocated to the platform's services, not your servers as a whole. There are no hosting costs — the platform runs on your infrastructure.

Pricing, licensing and getting started

See AI work with your business data.

In 30 minutes, we show how the assistant finds data, queries metrics and creates dashboards. Together, we discuss deployment in your data center, your choice of model and the right access permissions.

  • A live demo of the AI assistant and MCP tools
  • Model choice, data flows and governance for your use case
  • Next steps for a proof of concept in your environment