Ordered roughly by how much structure is preserved. Lower-on-the-stack approaches return more raw material; higher-up approaches return more direct answers.
π
Blob / Raw File Access
aka filesystem MCP, raw `git show`, file-read tools
Lowest fidelity
Hand the agent a file path or blob URL. It reads the bytes. No indexing, no understanding, no structure.
What it stores
Nothing. Just exposes the filesystem or blob storage as-is.
How agents query it
read_file(path), list_dir(path), cat repo/file.rb
Strengths
Always current. No index lag. Zero infra. Works on any repo.
Limitations
Agent burns context just exploring. No way to ask "where is X used?" without grepping the whole tree. Doesn't scale beyond ~hundreds of files.
Examples: Filesystem MCP, GitHub/GitLab raw API MCP servers, Cursor's basic file tools.
π
Lexical Search
aka grep, regex, full-text search, BM25
String-level
Index every token across every file. Match queries to text. Returns ranked file/line hits.
What it stores
Inverted index of tokens (and often regex-able n-grams) over the source corpus.
How agents query it
search("UserController"), grep -r "DELETE FROM users"
Strengths
Fast, precise for exact matches. Predictable. Great for "find this string." Mature tech.
Limitations
No understanding of synonyms, types, scope, or callers. "Find what calls this" requires guessing the right string. Misses paraphrased code.
Examples: Sourcegraph code search, GitLab Exact Code Search, ripgrep, classic IDE find-in-files.
π§¬
Semantic Search / RAG
aka vector embeddings, ANN retrieval, chunked context
Similarity-based
Chunk every file. Embed each chunk into a vector. Embed the query. Return the top-K nearest neighbors. Stuff them into the model's context window.
What it stores
A vector database of chunked code/docs. Each row: the chunk + its embedding + metadata.
How agents query it
retrieve("how does auth work?", k=10) β returns 10 vaguely-related code snippets.
Strengths
Handles fuzzy/natural-language queries. Good on prose, docs, README content. Surfaces "looks similar" matches.
Limitations
Code isn't prose. Chunks split functions mid-body. "Similar embedding" β "actually called." No precision for "what depends on this?" Stale the moment a file changes.
Examples: Most Cursor/Continue indexers, Glean (enterprise docs), Pinecone-backed RAG pipelines, Copilot Chat retrieval. What people mean by "RAG" usually = this.
βοΈ
Symbol Index / LSP
aka tree-sitter, ctags, language server, LSIF, SCIP
Structural, single-language
Parse source into ASTs. Resolve symbols (functions, classes, variables) and link definitions to references using a language server.
What it stores
Per-file ASTs and a symbol table linking definitions β references, scoped by language.
How agents query it
goto_definition(), find_references(), list_symbols(file)
Strengths
Compiler-accurate within a language. Knows real call sites, not string matches. The IDE experience.
Limitations
One language at a time. Stops at the repo boundary. Doesn't know about CI/CD, MRs, deployments, services, or org structure.
Examples: Sourcegraph precise code intel, GitNexus, Graphify, every IDE's go-to-definition, Moderne LSTs.
β
SDLC Knowledge Graph β Orbit / GKG
aka structured graph of code + CI/CD + MRs + deploys
Orbit's approach
Index code structure (functions, classes, calls) and SDLC entities (projects, MRs, pipelines, jobs, deployments, vulnerabilities) into a single typed graph. Agents traverse relationships across the whole engineering system.
What it stores
24 node types across 6 domains: core, source code, code review, CI, plan, security. Edges encode "calls", "defined in", "depends on", "deployed by", "fixes", etc.
How agents query it
query_graph(...) + get_graph_schema() via MCP or REST. Multi-hop graph traversal returns structured answers.
Strengths
Answers blast-radius, dependency, pipeline-inheritance, and cross-service questions in a single query. Cross-language. Native to GitLab, no ETL.
Limitations
Graph is only as rich as the schema. Less useful for natural-language "what is this concept" questions than RAG. Best paired with semantic/lexical for fuzzy intent β structured query.
Example query
// "What services depend on AuthService?"
MATCH (s:Service)
-[:CALLS]->(a:Service {name:
"AuthService"})
RETURN s.name, s.team
Returns
checkout-service // payments
notification-service // platform
admin-portal // internal-tools
π§
Agent Memory Graph
aka temporal knowledge graph, episodic memory
Conversation-scoped
Extract entities, relationships, and time-bounded facts from conversations and unstructured input. Persist them so an agent remembers what it learned across sessions.
What it stores
Entities + facts derived from chat, email, CRM, docs. Each fact has a validity window (from / to).
How agents query it
recall("what did the user say about pricing?") β returns ranked facts with timestamps.
Strengths
Solves the "agent forgets between sessions" problem. Tracks how facts change over time.
Limitations
No code structure, no SDLC awareness. Schema is general-purpose, you build the ingestion. Different problem domain entirely.
Examples: Graphiti, Zep, mem0. Complementary to Orbit, not competitive β an agent could use Zep for session memory and Orbit for SDLC queries simultaneously.