Docs RAG Assistant
A retrieval-augmented-generation (RAG) chat feature that answers user questions about a deployment using its own public user docs as the knowledge base. Both retrieval and generation route through the gateway itself.
Architecture
/v1/rag/chat is a thin orchestrator: it retrieves in-process, then calls the
gateway’s own public API as a user for the model work.
POST /v1/rag/chat (JWT-gated)
1. embed query ── HTTP ─► POST {RAG_API_BASE_URL}/embeddings (bge-m3)
2. cosine top-k over the committed JSON index (in-process)
3. build grounded prompt with citations (in-process)
4. generate ── HTTP ─► POST {RAG_API_BASE_URL}/chat/completions (qwen3.6-35b)
→ SSE: sources event, then the proxied OpenAI chunks, then [DONE]
Both HTTP calls carry RAG_API_KEY, so they flow through the standard
/v1/embeddings and /v1/chat/completions handlers → logged to api_logs and
counted toward cost / quota / concurrency. They also carry
X-On-Behalf-Of: <end-user id>, so that attribution lands on the real end
user (verified by JWT at /v1/rag/chat), not on the shared RAG_API_KEY account.
Corpus: the active distribution overlay’s documentation source (
<overlay>/content/docs/docs/source/*.md) — the same markdown that builds that deployment’s public doc site. A checkout with no overlay has no corpus; setRAG_CORPUS_DIRto your own documentation.Vector store: a plain JSON file (
<overlay>/content/rag/docs_index.json) scanned with pure-Python cosine similarity. The corpus is tiny, so no numpy / ANN index is needed. The index is committed (embeddings rounded to 6 decimals, ~0.9 MB) and lives inside theservingpackage so it ships in the Docker image — a fresh container serves retrieval immediately, with no build-time embedding call.Why call the gateway as a user (over HTTP) instead of the in-process router? So RAG requests are observable and metered. Direct
RouteExecutor/ adapter calls bypass the per-request logging, cost, quota, and concurrency that live in the/v1/*route handlers. Routing the model work back through those endpoints reuses all of it for free.
Code map
Path |
Role |
|---|---|
|
Heading-aware markdown chunking |
|
|
|
JSON vector store + cosine search |
|
Prompt assembly + sources payload |
|
|
|
|
|
Chat UI ( |
|
Streaming SSE client |
Rebuilding the index
The committed index is prebuilt with real bge-m3 embeddings. Regenerate it
(e.g. after the docs change) with the default gateway embedder:
RAG_GATEWAY_API_KEY=hyi-xxx make rag-ingest # real bge-m3, 1024-dim
RAG_GATEWAY_BASE_URL defaults to http://localhost:8080/v1 — embedding is
billable work, so a clone draws on its own gateway rather than on whoever wrote
the default. Point it at the gateway you want to embed through; the key must be
a valid user API key on that gateway. Chunks are embedded with the same
bge-m3 model the serving endpoint uses at query time, so query and document
vectors share one space.
For an offline run with no gateway/key (weak retrieval — dev/CI only):
RAG_EMBEDDER=hash make rag-ingest
The serving endpoint embeds the query with whichever embedder built the index (recorded in the index metadata) and fails loud (HTTP 502) if the query vector’s dimension doesn’t match the index. Always rebuild after switching embedders.
Deployment
The index ships inside the image (it lives under serving/), so no extra deploy
step is required. To refresh it, rebuild the image after re-running
make rag-ingest. /v1/rag/status reports index_loaded, the embedder mode,
and chunk count for a post-deploy check. If the embedding backend is unavailable
at query time, /v1/rag/chat returns a graceful 503 rather than a 500.
Required env per deployment: set RAG_API_KEY to a valid user API key, and
point RAG_API_BASE_URL at the gateway’s own address for that environment — the
default http://localhost:8080/v1 matches the prod Docker container’s port, but
staging (systemd) binds 8000, so it needs RAG_API_BASE_URL=http://localhost:8000/v1.
The inner calls present RAG_API_KEY as the credential but carry
X-On-Behalf-Of: <end-user id>, so cost / quota / logs / per-user concurrency
attribute to the real end user (verified by JWT at /v1/rag/chat) rather than
to the shared service account — each user’s RAG usage counts against their own
daily quota. verify_api_key honors X-On-Behalf-Of only for the configured
RAG_API_KEY; any other key’s header is ignored, and if RAG_API_KEY is unset
impersonation is disabled entirely.
Endpoints
Both live under the gateway and authenticate with the dashboard JWT
(get_current_user), so the Next.js chat page calls them with the session token
it already holds.
GET /v1/rag/status— whether the index is built, chunk count, models.POST /v1/rag/chat— body{ messages, top_k?, stream? }. The generation model is fixed server-side (RAG_CHAT_MODEL); it is not client-selectable, so the endpoint can’t be used to reach role-gated models.Streaming (default): SSE — first a
{"type":"sources", ...}event, then OpenAI-format completion chunks, then[DONE].Non-streaming:
{ answer, sources, model }.
Logging & quota
The RAG model calls go through the gateway’s own /v1/embeddings and
/v1/chat/completions, so they land in api_logs and count toward cost, daily
quota, and per-user concurrency — attributed to the real end user via the
X-On-Behalf-Of header (the JWT-verified caller of /v1/rag/chat), with
RAG_API_KEY as the presented credential. A row is identifiable as RAG-originated
by its metadata.user_agent = "doc_assistant". Set RAG_API_KEY to a valid user
API key; when it is unset the endpoint returns 503. An upstream 429
(quota/rate) is passed through to the caller — note this can now be the end
user’s own daily quota, not the service account’s.
curl -sN https://<your-gateway>/v1/rag/chat \
-H "Authorization: Bearer <jwt>" -H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"How do I get an API key?"}]}'
Configuration
All optional; sensible defaults resolve relative to the repo root.
Env var |
Default |
Purpose |
|---|---|---|
|
(unset) |
User API key the handler calls the gateway with (required at serving time) |
|
|
Gateway the handler calls (self-call for logging/quota) |
|
the overlay’s |
Vector index location |
|
the overlay’s |
Markdown corpus |
|
|
|
|
|
Gateway used by ingest (gateway mode) |
|
|
Embedding model id (gateway mode) |
|
|
Answer-generation model |
|
|
Chunks retrieved per query |
|
|
Answer token budget |
|
|
Generation temperature |
Prototype limitations
The committed index is refreshed manually (
make rag-ingest+ rebuild image), not on a schedule — it can lag the docs until regenerated.The
HashEmbedderfallback exists only so the pipeline runs without the gateway (dev/CI); its retrieval quality is weak.No answer caching, no reranking, and history is truncated to the last few turns. The store is loaded into memory per process (cached, mtime-invalidated).