Skip to content
All work
AI InfrastructureOpen source · local-first

AxiomAI

A production control plane for AI applications.

One gateway between your applications and OpenAI, Gemini and Anthropic — with routing, failover, cost tracking and grounded retrieval built in.

Problem

Applications that call LLM vendors directly inherit every vendor's outages, pricing and SDK quirks. Keys leak into services, costs are invisible until the invoice arrives, and there is no single place to enforce limits or ground answers in company knowledge.

Context

Built as a systems project to demonstrate production AI engineering rather than another chatbot UI. Local-first by design: Docker Compose for Postgres + pgvector, a FastAPI modular monolith, and a Next.js operator dashboard.

Architecture

Applications never talk to model vendors directly. Every request passes a linear, individually testable pipeline before a deterministic router chooses a provider adapter.

AxiomAI architecture
  1. Application
    • Client app
    • SDK / HTTP
    • Operator dashboard
  2. AxiomAI Gateway
    • API-key auth
    • Scopes
    • Validation
    • Rate limit + quota
  3. Provider Router
    • Deterministic routing
    • Timeout + retry
    • Fallback chain
    • Circuit breaker
  4. Providers
    • OpenAI
    • Gemini
    • Anthropic
    • FakeProvider
  5. Persistence
    • PostgreSQL 16
    • pgvector
    • Usage + cost
    • Attempt log

Engineering decisions

Modular monolith, not microservices
One engineer can run, test and demo the whole system locally, and requests, usage and cost stay transactionally consistent. Services split only when a real scaling or team boundary appears.
Provider adapters behind one protocol
Routing, cost and auth code never import a vendor SDK, so adding or removing a provider does not touch business logic.
Deterministic routing before any ML router
Predictable, debuggable behaviour comes first. Fallback only runs for errors where retrying elsewhere makes sense — never for invalid requests or content filters.
Pricing as data
Model prices live in a model_pricing table rather than hard-coded rates, so cost per request stays correct when vendors change prices.
Fake providers for tests
80+ pytest cases run against fake chat and embedding providers, so CI never spends live credits and failure paths can be reproduced on demand.

Reliability

Each provider call is wrapped in a timeout with selective, exponentially backed-off retries. When a vendor keeps failing, a circuit breaker stops sending it traffic and the fallback chain moves to the next provider. Every hop is recorded, so an operator can see exactly why a request ended where it did.

  • Timeout
  • Selective retry
  • Exponential backoff
  • Fallback chain
  • Circuit breaker

Security

Application keys are high-entropy, Argon2-hashed and shown in plaintext exactly once. Scopes are deny-by-default, provider secrets live only in server environment variables, logs never contain keys or prompts, and every query is scoped to the calling application — with isolation covered by tests.

  • Argon2-hashed keys
  • Deny-by-default scopes
  • Env-only secrets
  • Application isolation

RAG

Knowledge is ingested as size-limited UTF-8 text or Markdown, chunked, embedded with Gemini and stored in pgvector. A constrained retrieval agent has one fixed action, bounded context and no tool loop. Retrieved text is treated as untrusted data, and sources are returned as structured data independent of what the model writes.

  • Ingest
  • Chunk
  • Embed
  • Retrieve
  • Cite

Observability

Each request gets a request id and trace id. Usage, cost, provider attempts, retries and failure reasons are persisted so the dashboard reads real data rather than mock charts.

Lessons

  • Deciding what is out of scope — autonomous agents, evaluation, Kubernetes — kept the core reliable.
  • Treating retrieved content as untrusted is a design constraint, not an afterthought.
  • Fake providers make failure paths testable; live vendors make them expensive.