This page uses diagrams and tables designed for larger screens. Open it on your laptop for the full experience.
A full Orbit deployment running inside your network. Your SDLC data never leaves your environment. GKG, Siphon, NATS, and ClickHouse — all on Kubernetes, connected to your existing GitLab instance.
Choose based on how you run GitLab today. Both paths deliver the same GKG capabilities.
Your GitLab instance runs on a Linux server via the Omnibus package (the standard self-managed install). GKG runs on a separate Kubernetes cluster that you provision alongside it. The two communicate over TLS and don't need to share a network.
gitlab-ee or gitlab-ce on a VM or bare metal.Your GitLab instance already runs on Kubernetes using the official gitlab/gitlab Helm chart. You extend your existing cluster by deploying the GKG Helm charts alongside it - no new cluster needed, GKG connects to GitLab via internal Kubernetes services.
helm install gitlab gitlab/gitlab.gitlab/gitlab chart. No changes to your GitLab pods - GKG components run alongside in the same cluster and connect via internal Kubernetes services.
Four components. Three of them are new. One (ClickHouse) you provision separately.
The indexing and query engine. Processes raw SDLC events from Siphon, builds the graph, and serves the query_graph and get_graph_schema API endpoints. Runs as two separate workloads: Indexer (write-heavy) and Web (query-serving).
Captures PostgreSQL WAL (write-ahead log) from your GitLab database and streams SDLC events into NATS. Siphon is the bridge between your existing GitLab data and the knowledge graph. It reads from PostgreSQL logical replication slots - no schema changes to your GitLab DB.
High-performance messaging bus that buffers events between Siphon and the GKG indexer. Runs as a 3-node StatefulSet for fault tolerance. Deployed and managed by the GKG Helm chart - you don't configure NATS directly beyond sizing.
The analytics database that stores the knowledge graph. You provision this separately - either ClickHouse Cloud or self-managed ClickHouse on your own infrastructure. GKG reads and writes to it; you own the sizing, backups, and upgrades. Baseline storage: roughly one-third of your total PostgreSQL database size.
GitLab engineering works with you hands-on for all design partner deployments. This is what the process looks like end to end.
Stand up a dedicated K8s cluster (GKE, EKS, AKS, or on-prem) for GKG. It does not need to share a VPC with your GitLab Omnibus server - just reachability over TLS.
Provision a 3-node ClickHouse cluster (ClickHouse Cloud or self-hosted) and note the connection string. Storage baseline is roughly one-third your PostgreSQL database size.
Siphon reads from PostgreSQL WAL. You need to enable logical replication and create a replication slot. No schema changes to your GitLab DB required.
Add the GitLab Helm repository and install the GKG chart with your configuration. GitLab will provide the chart values file as part of the onboarding.
gkg-values.yaml during onboarding. It includes your ClickHouse credentials, PostgreSQL replication config, and GKG endpoint settings.In your GitLab instance admin panel, configure the GKG endpoint URL and authentication token. This links your GitLab instance to the GKG web service running on your K8s cluster.
As a GitLab admin or top-level group Owner, enable Orbit indexing per group. Indexing starts on the main branch. Full initial index time depends on repo count and size.
This path is for customers already running gitlab/gitlab on Kubernetes. If you're on Omnibus, use the OAK path instead.
Same as OAK: provision a 3-node ClickHouse cluster separately. ClickHouse Cloud or self-hosted both work. Storage baseline is roughly one-third your PostgreSQL DB size.
Install the GKG chart into the same namespace as your GitLab deployment. GKG will discover your PostgreSQL and Gitaly services automatically via Kubernetes DNS.
gitlab release lets Siphon connect to PostgreSQL and Gitaly via internal Kubernetes service discovery - no external networking required.Same final steps as OAK: configure the GKG endpoint in admin settings, then enable indexing per group from admin or group Owner settings.
These are baseline specs. Production sizing will vary based on group count, repo sizes, and query volume.
| Component | Nodes | CPU / node | RAM / node | Storage | Notes |
|---|---|---|---|---|---|
| ClickHouse | 3 | ~4 vCPU | ~16 GB | ~1/3 of PostgreSQL DB size | Customer-provisioned. Not bundled. Cloud or self-hosted. |
| NATS | 3 | ~4 vCPU | ~8 GB | Minimal (messaging buffer) | Deployed by GKG chart as StatefulSet. 3-node HA. |
| GKG Indexer | 2+ (HA) | ~4 vCPU | ~3 GB+ | Ephemeral (stateless) | 3+ GB per pod for large or complex repositories. |
| GKG Web | 2+ (HA) | ~2 vCPU | ~1 GB | Ephemeral (stateless) | Query-serving pods. Minimum 2 for HA. |
| Siphon | 1+ | ~2 vCPU | ~2 GB | Ephemeral (stateless) | Single instance acceptable; HA config available. |
| Persistent Volumes | - | - | NATS: ~20 GB per node | PVCs for NATS StatefulSet. StorageClass must support ReadWriteOnce. | |
Go through this list before your onboarding call with the GitLab team.
wal_level = logical)Indexing is controlled by admins at the group level. Here's what to plan for.
We're onboarding a small set of self-managed design partners now. You'll work directly with the GKG engineering team and shape the deployment experience before GA.
Join the Self-Managed WaitlistYou'll hear back from the GKG team within a few business days.