Papers
Topics
Authors
Recent
Search
2000 character limit reached

OGX: Secure Multi-Tenant RAG System

Updated 5 July 2026
  • OGX is an open-source framework that integrates attribute-based access control with retrieval-augmented generation for secure, multi-tenant enterprise AI.
  • It decouples the API from inference providers, enabling vendor neutrality and supporting pluggable vector stores and models.
  • Empirical evaluations demonstrate that gated retrieval eliminates cross-tenant leakage and improves precision with minimal latency overhead.

Searching arXiv for the cited OGX paper and closely related terms to ground the article in current arXiv metadata. OGX is the concrete system implementation of a security architecture for enterprise Retrieval-Augmented Generation and agentic AI. Presented as an open-source, vendor-neutral framework—formerly Llama Stack, now “Open GenAI Stack”—it exposes OpenAI-compatible APIs while enforcing multitenant isolation, retrieval authorization, and server-side orchestration. Its central premise is that, in enterprise multitenant RAG, relevance is not authorization: standard dense, sparse, and hybrid retrieval optimize similarity or keyword overlap rather than policy, so a query from one tenant can surface another tenant’s confidential data if that content ranks highest. OGX addresses that mismatch by inserting authorization checks directly into retrieval and orchestration paths, so only authorized content reaches model context (Arceo et al., 6 May 2026).

1. Concept and system scope

OGX is introduced for enterprise settings in which multiple tenants share infrastructure, data has heterogeneous sensitivity, fine-grained authorization is mandatory, and organizations require portability across model vendors and deployment environments. In that setting, the paper argues that architectures designed for single-tenant or consumer use break down because retrieval is typically optimized for relevance and orchestration is often delegated to the client. OGX is designed to solve that mismatch without requiring full per-tenant duplication of the entire stack (Arceo et al., 6 May 2026).

The framework is vendor-neutral because its API layer is decoupled from the provider layer. Each API can be backed by pluggable providers, which may be inline or remote, and a routing layer resolves logical resources to concrete providers after incorporating tenant identity and authorization. This allows one tenant’s vector store or model to map to one backend and another tenant’s to a different backend, while the client continues to use the same API surface.

It is multitenant because tenant identity and attributes are part of the API and storage model from the beginning. The paper states that Responses, vector stores, files, conversations, prompts, and other entities are tenant-scoped API objects subject to the same declarative attribute-based access control semantics. This suggests that OGX is intended not merely as a retrieval wrapper, but as a unified security substrate for agentic resources.

2. Security model and the relevance-authorization gap

The paper identifies four major gaps in conventional enterprise agent deployments. The first is the mismatch between relevance-based retrieval and authorization: enterprise corpora are shared across tenants or business units, but retrieval engines optimize ranking, not policy. The second is tool-mediated disclosure, in which tools execute using agent credentials or server capabilities rather than the end user’s authorization scope. The third is context accumulation across turns: previously retrieved unauthorized material can persist through retained history, tool outputs, summaries, and cached intermediate state. The fourth is client-side orchestration bypass: a malicious or buggy client can skip gated retrieval, call ungated endpoints, manipulate tool invocations, or replay unauthorized outputs (Arceo et al., 6 May 2026).

The threat model is explicit. In scope are a malicious tenant trying to retrieve another tenant’s data, including by prompt injection; a compromised client bypassing authorization by calling ungated endpoints or manipulating tool invocations; and a buggy tool leaking cross-tenant state through outputs or side effects. Out of scope are a compromised server process, insider operators with infrastructure access, side-channel attacks, and model extraction. The server process is the trust boundary.

The system goals are stated as follows: G1: No cross-tenant data leakage through retrieval; G2: Authorization enforcement independent of orchestration mode; G3: Denied queries fail fast without invoking inference. The architecture also assumes A1: correct token-to-tenant mapping by the authentication provider; A2: immutable document ownership metadata assigned at ingestion; A3: the inference layer is untrusted—it may leak any context it receives, so isolation must happen before context is built; and A4: vector backends correctly apply metadata filters when predicate pushdown is supported.

The agent execution pattern is formalized as

E=[(i1,ϕ1,r1),(i2,ϕ2,r2),…,(in,∅,rn)].E = [(i_1, \phi_1, r_1), (i_2, \phi_2, r_2), \dots, (i_n, \emptyset, r_n)].

An agent run is therefore modeled as alternating inference and tool steps, with the final step having no tool call. The key retrieval security requirement is

{d∈D:relevance⁡(q,d)>θ∧P(u,d)=permit}.\{ d \in D : \operatorname{relevance}(q,d) > \theta \wedge P(u,d) = \mathrm{permit} \}.

The design principle is that retrieved documents must satisfy relevance and authorization simultaneously. A common misconception addressed by the paper is that strong ranking alone, or even relevance thresholds, are sufficient; OGX treats that view as the root cause of cross-tenant leakage.

3. Layered isolation architecture

The OGX architecture has two major halves: a layered isolation data path and a server-side control path. On the data path, Layer 1 is policy-aware ingestion. Every document is tagged with tenant ownership and access metadata at ingestion time, denoted as

I(d,t)→Dt.\mathcal{I}(d, t) \rightarrow D_t.

Each chunk inherits tenant and policy attributes early so downstream retrieval can filter consistently. The paper emphasizes ingestion-time tagging rather than retrofitting metadata later, because inconsistent metadata is described as a common source of accidental cross-tenant coupling (Arceo et al., 6 May 2026).

Layer 2 is retrieval gating, described as the main security mechanism. It consists of resource-level ABAC authorization before search and chunk-level filtering after retrieval. If the vector backend supports predicate pushdown, tenant or attribute filters are applied inside the search engine itself. If the backend does not support pushdown, OGX still enforces post-retrieval filtering, which preserves security but may hurt recall at scale because unauthorized chunks can occupy top-kk positions before filtering. The paper specifically recommends pushdown-capable backends such as pgvector, Qdrant, and Milvus for production.

Layer 3 is shared inference. Once the first two layers ensure that only authorized content enters the context window, the model endpoint itself can be shared across tenants. The cost reduction is described as going from per-tenant duplication O(N⋅M)O(N \cdot M) to shared serving O(M)O(M), where NN is the number of tenants and MM the number of model endpoints. The limitation is explicit: this does not solve parametric memory leakage from pretraining.

Surrounding these layers is server-side orchestration, which secures the control path. The orchestration flow is: Request → Input safety → Inference → Tool execution → Output safety → Response storage. The paper explicitly separates the roles of gating and orchestration: gating provides the data isolation guarantee, while server-side orchestration prevents the client from bypassing it. In the paper’s formulation, gating gives the security guarantee, and server-side orchestration gives the enforcement guarantee.

4. API model, provider abstraction, and tenant-scoped resources

OGX exposes more than 20 APIs spanning inference, agents/Responses API, vector stores, search, safety, tools, telemetry, files, file processors, prompts, conversations, connectors, and tool groups. The centerpiece is the Responses API, which implements the OpenAI Responses API paradigm as an open-source server endpoint. In the paper’s usage, “OpenAI-compatible” refers to interface compatibility, while “open-source Responses API” refers to an OSS, self-hostable implementation with transparent, pluggable backends and server-enforced multitenancy (Arceo et al., 6 May 2026).

The supported providers listed in the appendix include inference backends such as vLLM, Ollama, OpenAI, Anthropic, Azure, AWS Bedrock, Databricks, Gemini, Together, NVIDIA, and WatsonX; and vector stores such as Chroma, pgvector, Elasticsearch, Qdrant, Weaviate, Milvus, OCI, FAISS, and sqlite-vec. This provider indirection is one of the main reasons the paper describes OGX as vendor-neutral.

Authorization is declarative ABAC. The default policy permits access if the user owns the resource or if the user’s attributes—roles, teams, projects, namespaces—match the resource’s access attributes, and default deny applies otherwise. Enforcement occurs in three places: API route middleware, routing table resolution before logical resources are mapped to providers, and storage read time where user attributes are translated into query filters. External identity from JWT or Kubernetes auth is mapped into these attributes so application developers do not need to hard-code tenant IDs into business logic.

The detailed API interaction described in the appendix makes the Responses API the orchestration hub. A client can POST to /v1/responses, and the server-side Responses API may read prompts, read or write conversation state, trigger compaction, call vector store and search APIs, read files, invoke file processors, and route tool calls including file search, web search, image generation, code interpreter, shell, connectors, and remote MCP. The Conversations API plus Compaction API supports summarization of long conversations when token limits are exceeded, while preserving user messages verbatim. The paper does not provide a formal proof about compaction safety, but the intended architecture is that both original and compacted state remain tenant-scoped and subject to per-turn authorization.

The integration model with client-side frameworks is hybrid rather than exclusionary. LangChain, LangGraph, CrewAI, and similar frameworks can continue to compose agent logic on the client, but call into OGX’s standard endpoints for inference, retrieval, and orchestration. The caveat is also explicit: client-side function tools execute outside the server trust boundary, so OGX can classify them but cannot enforce server-side invariants on client-executed code.

5. Empirical evaluation

The evaluation is organized around a 2×2 matrix: Config A: client-side orchestration, ungated retrieval; Config B: client-side orchestration, gated retrieval; Config C: server-side orchestration, ungated retrieval; Config D: server-side orchestration, gated retrieval. The setup uses three synthetic tenants—finance, engineering, and legal—with 300 total documents, 100 per tenant, around 512 tokens each, and controlled topical overlap. For each configuration, the paper reports 300 authorized queries, 300 cross-tenant probes, and 90 prompt injection probes across instruction override, role impersonation, debug exploitation, and context manipulation (Arceo et al., 6 May 2026).

The primary software stack uses OpenAI gpt-4o-mini as the model via OGX’s remote provider, OpenAI text-embedding-3-small for embeddings, sqlite-vec as the vector store, and a mock auth layer mapping bearer tokens to tenant identities. GPU infrastructure measurements use vLLM on an NVIDIA T4 serving Llama-3.2-1B-Instruct. The central security metrics are Cross-Tenant Leakage Rate, defined as the fraction of cross-tenant probes returning at least one unauthorized chunk, and Authorization Violation Rate, defined as the fraction of all API calls returning unauthorized data.

The headline result is that ABAC gating eliminates cross-tenant leakage. Under assumptions A1–A4, the architecture achieves CTLR = 0 and AVR = 0 for any query workload. Ungated retrieval, by contrast, leaked cross-tenant data in 98–100% of probes; elsewhere the paper summarizes the same phenomenon as “nearly all cross-tenant probes” leaking regardless of orchestration mode. The stronger conclusion is that orchestration mode by itself is not the security control. Without gating, both client-side and server-side orchestration are vulnerable. With gating, both can block leakage, but only server-side orchestration prevents bypass by a malicious client.

The prompt injection experiments reinforce that conclusion. Under gated configurations, all 90 prompt injection probes caused 0% leakage. Under ungated configurations, 62–80% of probes retrieved cross-tenant data. The paper interprets this as a retrieval-layer defense rather than a prompt-following defense: no matter how malicious the prompt, unauthorized documents are never returned by the vector store under ABAC gating.

The evaluation also reports retrieval-quality effects. Using synthetic embeddings with about 0.95 cross-tenant similarity to stress retrieval, ungated retrieval leaked 52% of queries. Chunk-level gating improved Precision@5 from 0.200 to 0.433, a 2.2× improvement, increased MRR from 0.700 to 1.000, and left Recall@5 at 1.000 in that controlled setting. Per-tenant indexing achieved the same Recall@5 and MRR but lower Precision@5 than chunk-level gating in that benchmark. This suggests that, in shared corpora, gating can improve retrieval quality by filtering cross-tenant noise.

Scaling behavior depends heavily on predicate pushdown. With sqlite-vec, which lacks native pushdown, OGX performs post-retrieval filtering. CTLR remains at 0%, but Recall@5 declines sharply with corpus growth: at 100 chunks, Recall@5 = 1.000; at 1,000 chunks, 0.100; at 10,000 chunks, 0.010; at 50,000 chunks, 0.002. Filter overhead remains small, from 0.74 ms to 2.95 ms. This is the basis for the paper’s production recommendation to use pushdown-capable vector stores.

Latency overhead is characterized as negligible relative to inference. Isolating search from inference, gated search adds about 19 ms total: auth server round-trip of approximately 14 ms, ABAC policy evaluation of less than 1 ms, and per-tenant store lookup of approximately 5 ms. In the GPU appendix, vLLM direct baseline is 447.9 ms median and OGX routing plus dispatch is 452.6 ms median, so routing adds 4.7 ms, or about 1.0% of baseline inference. Search latency rises from 283.9 ms ungated to 289.4 ms tenant-gated, so metadata filtering adds 5.5 ms, or about 1.9% of search time. End-to-end authorized-query latency is dominated by external API variability: Config A mean 4208 ms, p50 3600 ms, p99 10818 ms; Config B mean 3851 ms, p50 3427 ms, p99 9795 ms; Config C mean 7620 ms, p50 7507 ms, p99 16462 ms; Config D mean 6934 ms, p50 6431 ms, p99 14623 ms. The consistent pattern is that server-side orchestration adds roughly 3 seconds in these non-streaming experiments because the Responses API executes the multi-turn tool loop server-side before returning the full result. The paper notes that streaming is supported and would materially reduce time-to-first-token.

Throughput scales linearly up to concurrency c=25c = 25, and client-side orchestration has roughly 2× the throughput of server-side orchestration at high concurrency because the request path is shorter. At c=25c = 25, Config A reaches 5.4 QPS, Config B 4.2 QPS, Config C 2.2 QPS, and Config D 2.6 QPS. The appendix also mentions a 48-case ABAC correctness matrix with 100% accuracy and 0% false positives.

6. Tradeoffs, deployment patterns, and broader significance

The paper identifies several design tradeoffs. ABAC is expressive, but policy complexity grows with the number of tenants, roles, and resource types, and deeply nested permission hierarchies may create operational overhead. Predicate pushdown is a second tradeoff: without backend support for metadata predicate pushdown, security still holds but recall suffers badly at scale. OGX therefore separates security guarantees from retrieval-quality guarantees. A third tradeoff is that client-side function tools remain a weak point by design, since execution outside the server trust boundary limits enforcement. A fourth limitation is model memory: OGX constrains retrieved and orchestrated context, but does not claim to prevent a model from emitting information memorized during pretraining or prior fine-tuning (Arceo et al., 6 May 2026).

The architecture is also presented as unnecessary in some cases. The paper states that it is overkill for single-tenant deployments, public data, organizations that already isolate each tenant by separate infrastructure, or teams that prefer fully managed SaaS over self-hosted infrastructure. If all tenants are trusted and latency sensitivity dominates, client-side orchestration may be a better fit.

Deployment guidance centers on Kubernetes and provider portability. The OGX Kubernetes Operator supports shared multitenant instances with ABAC isolation, per-tenant instances with namespace-level Kubernetes RBAC, or hybrids. The same provider abstraction allows migration from lightweight local backends such as sqlite-vec or in-process safety models to production systems such as pgvector or vLLM without changing agent code or security policy definitions. The paper cites the main OGX repository and the OGX Kubernetes Operator repository as open-source artifacts.

In broader terms, OGX is presented not as a brand-new primitive, but as a composition of ABAC, metadata-aware retrieval, shared serving, and server-side orchestration tailored to enterprise agentic AI. The novelty lies in formalizing the “relevance-authorization gap,” identifying where consumer-style assumptions fail under multitenancy, and showing empirically that retrieval-time ABAC gating can eliminate cross-tenant leakage without meaningful performance cost. A plausible implication is that OGX’s main contribution is architectural: it makes shared inference economically viable in multitenant deployments while enforcing authorization at the point where retrieved and tool-produced content enters model context.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to OGX.