Context Architecture

How AI agents actually get code context

Six fundamentally different approaches to giving an AI agent visibility into a codebase and SDLC. Each makes a different tradeoff between recall, precision, freshness, and the kind of question it can answer. Most tools combine two or three. GKG is one of them.

Six Approaches
From raw bytes to structured graph
Ordered roughly by how much structure is preserved. Lower-on-the-stack approaches return more raw material; higher-up approaches return more direct answers.
πŸ“„
Blob / Raw File Access
aka filesystem MCP, raw `git show`, file-read tools
Lowest fidelity
Hand the agent a file path or blob URL. It reads the bytes. No indexing, no understanding, no structure.
What it stores
Nothing. Just exposes the filesystem or blob storage as-is.
How agents query it
read_file(path), list_dir(path), cat repo/file.rb
Strengths
Always current. No index lag. Zero infra. Works on any repo.
Limitations
Agent burns context just exploring. No way to ask "where is X used?" without grepping the whole tree. Doesn't scale beyond ~hundreds of files.
Examples: Filesystem MCP, GitHub/GitLab raw API MCP servers, Cursor's basic file tools.
πŸ”
Lexical Search
aka grep, regex, full-text search, BM25
String-level
Index every token across every file. Match queries to text. Returns ranked file/line hits.
What it stores
Inverted index of tokens (and often regex-able n-grams) over the source corpus.
How agents query it
search("UserController"), grep -r "DELETE FROM users"
Strengths
Fast, precise for exact matches. Predictable. Great for "find this string." Mature tech.
Limitations
No understanding of synonyms, types, scope, or callers. "Find what calls this" requires guessing the right string. Misses paraphrased code.
Examples: Sourcegraph code search, GitLab Exact Code Search, ripgrep, classic IDE find-in-files.
🧬
Semantic Search / RAG
aka vector embeddings, ANN retrieval, chunked context
Similarity-based
Chunk every file. Embed each chunk into a vector. Embed the query. Return the top-K nearest neighbors. Stuff them into the model's context window.
What it stores
A vector database of chunked code/docs. Each row: the chunk + its embedding + metadata.
How agents query it
retrieve("how does auth work?", k=10) β†’ returns 10 vaguely-related code snippets.
Strengths
Handles fuzzy/natural-language queries. Good on prose, docs, README content. Surfaces "looks similar" matches.
Limitations
Code isn't prose. Chunks split functions mid-body. "Similar embedding" β‰  "actually called." No precision for "what depends on this?" Stale the moment a file changes.
Examples: Most Cursor/Continue indexers, Glean (enterprise docs), Pinecone-backed RAG pipelines, Copilot Chat retrieval. What people mean by "RAG" usually = this.
βš™οΈ
Symbol Index / LSP
aka tree-sitter, ctags, language server, LSIF, SCIP
Structural, single-language
Parse source into ASTs. Resolve symbols (functions, classes, variables) and link definitions to references using a language server.
What it stores
Per-file ASTs and a symbol table linking definitions ↔ references, scoped by language.
How agents query it
goto_definition(), find_references(), list_symbols(file)
Strengths
Compiler-accurate within a language. Knows real call sites, not string matches. The IDE experience.
Limitations
One language at a time. Stops at the repo boundary. Doesn't know about CI/CD, MRs, deployments, services, or org structure.
Examples: Sourcegraph precise code intel, GitNexus, Graphify, every IDE's go-to-definition, Moderne LSTs.
β˜…
SDLC Knowledge Graph β€” Orbit / GKG
aka structured graph of code + CI/CD + MRs + deploys
Orbit's approach
Index code structure (functions, classes, calls) and SDLC entities (projects, MRs, pipelines, jobs, deployments, vulnerabilities) into a single typed graph. Agents traverse relationships across the whole engineering system.
What it stores
24 node types across 6 domains: core, source code, code review, CI, plan, security. Edges encode "calls", "defined in", "depends on", "deployed by", "fixes", etc.
How agents query it
query_graph(...) + get_graph_schema() via MCP or REST. Multi-hop graph traversal returns structured answers.
Strengths
Answers blast-radius, dependency, pipeline-inheritance, and cross-service questions in a single query. Cross-language. Native to GitLab, no ETL.
Limitations
Graph is only as rich as the schema. Less useful for natural-language "what is this concept" questions than RAG. Best paired with semantic/lexical for fuzzy intent β†’ structured query.
Example query
// "What services depend on AuthService?"
MATCH (s:Service)-[:CALLS]->(a:Service {name:"AuthService"})
RETURN s.name, s.team
Returns
checkout-service // payments
notification-service // platform
admin-portal // internal-tools
🧠
Agent Memory Graph
aka temporal knowledge graph, episodic memory
Conversation-scoped
Extract entities, relationships, and time-bounded facts from conversations and unstructured input. Persist them so an agent remembers what it learned across sessions.
What it stores
Entities + facts derived from chat, email, CRM, docs. Each fact has a validity window (from / to).
How agents query it
recall("what did the user say about pricing?") β†’ returns ranked facts with timestamps.
Strengths
Solves the "agent forgets between sessions" problem. Tracks how facts change over time.
Limitations
No code structure, no SDLC awareness. Schema is general-purpose, you build the ingestion. Different problem domain entirely.
Examples: Graphiti, Zep, mem0. Complementary to Orbit, not competitive β€” an agent could use Zep for session memory and Orbit for SDLC queries simultaneously.
Side by Side
When each approach wins
No approach is best for every question. The right answer is usually a stack: lexical for exact strings, semantic for fuzzy intent, GKG for structured SDLC traversal.
Approach Best For Cross-File Refs SDLC / CI Aware Fuzzy Queries Freshness Token Cost
Blob / Raw Reading a known file ● No ● No ● No ● Live ● High
Lexical Find exact string / regex ● Token only ● No ● No ● Near-live ● Medium
Semantic / RAG Natural-language intent ● No ● No ● Yes ● Reindex lag ● Medium
Symbol / LSP Goto-def, find refs ● Yes (per language) ● No ● No ● On-save ● Low
SDLC GraphOrbit Blast radius, deps, pipelines, MR↔deploy ● Yes (multi-lang) ● Yes ● Via schema ● Continuous ● Low
Memory Graph Cross-session agent memory ● No ● No ● Yes ● Live ● Low

Where Orbit fits in the stack

Most code-context products live in one row of that table. Cursor's indexer is semantic. Sourcegraph is lexical + symbol. Glean is semantic over docs. Graphiti is memory.

Orbit is the only one indexing the full SDLC as a typed graph β€” code structure plus pipelines, MRs, deployments, vulnerabilities, and group hierarchy β€” and exposing it as a first-class API any agent can query. That makes it the right tool when the question is structural ("what depends on X?", "which pipelines inherit this config?", "what changed in the deploy that broke prod?") rather than lexical or fuzzy.

For natural-language intent or exact string lookup, pair Orbit with semantic and lexical search. They're not the same tool. They're not solving the same problem.