Service / AI Infrastructure & Architecture

The stack
under everything.

AI isn't just software — it's infrastructure. We design and build the hardware, routing, and deployment architecture that makes production AI reliable, sovereign, and fast.

What we build Talk to us
The reality most people skip

A great model on
bad infrastructure is just
an expensive mistake.

Most AI deployments fail at the infrastructure layer, not the model layer. Wrong hardware, cloud dependencies that create data sovereignty problems, no fallback when a provider goes down, no routing logic between models. We build the foundation that makes everything else work.

The core problem

Cloud-only is a liability

When your AI runs entirely on third-party infrastructure, you have no control over latency, cost, data sovereignty, or availability. Every inference call is a dependency on someone else's uptime, pricing, and data handling decisions.

The alternative

Local-first, cloud-fallback

Run capable local models for the work they can handle — summarisation, classification, routine reasoning. Route complex tasks to cloud models only when needed. Your data stays local by default. Costs drop significantly. Latency is predictable.

What makes it hard

Getting it right takes experience

GPU configuration, model quantisation, routing logic, fallback handling, streaming, load balancing — each piece has failure modes that only surface under production load. We've made these mistakes so you don't have to.

What we bring

Built from bare metal up

We run our own multi-node inference cluster — dual RTX 5060 Ti nodes on Ubuntu bare metal, serving Qwen 30B models via llama.cpp at 52 tokens/second. We didn't read about this setup, we built it. That's what we bring to your infrastructure.

What Firebird delivers

End-to-end infrastructure
design and deployment.

From scoping your requirements to running production inference, we architect and build the full stack.

01

Inference Infrastructure

Local GPU cluster design and deployment. Model selection, quantisation, and serving configuration. We spec the hardware, configure the stack, and get you to production-ready local inference — llama.cpp, Caddy reverse proxy, SSL, streaming via Server-Sent Events.

02

Multi-Tier LLM Routing

Not every task needs your most powerful model. We build routing layers that direct requests to the right model for the job — local models for routine work, cloud routing (Anthropic, OpenRouter) for complex reasoning — with automatic fallback and model-agnostic application code.

03

Agent Orchestration

Infrastructure for AI agents that run autonomously — task queuing, failure detection, context management, and cognitive continuity across sessions. Built for production, not demos: four-mode failure detection, duplicate read protection, graduated context management.

04

Data Architecture

PostgreSQL with pgvector for hybrid semantic search. Multi-tenant isolation via EF Core global query filters. Soft delete across all entities. Single source of truth across all services. We design data architecture that serves both your application and your AI.

Our own stack

We run what
we recommend.

The infrastructure we deploy for clients is the same architecture we run ourselves. These aren't theoretical recommendations.

Inference nodes2x Ubuntu bare metal, NVIDIA RTX 5060 Ti (dual per node)
Total VRAM64GB across inference cluster
Primary modelQwen3-Coder-30B — 52 t/s, 1.78s latency
Serving layerllama.cpp + liteLLM
Cloud routingAnthropic Claude, OpenRouter (no OpenAI)
Reverse proxyCaddy with automatic SSL
OrchestrationDocker / docker-compose
DatabasePostgreSQL with pgvector
Why it matters

Sovereign infrastructure
isn't optional.

The conversation about AI and data sovereignty is happening at the boardroom level now — and rightly so. When your AI runs on cloud infrastructure you don't control, your customer data, your production records, and your institutional knowledge flow through third-party systems governed by laws that may not match your obligations.

We architect for sovereignty by default. Local inference means your data doesn't leave your network to power your AI. Cloud routing is available as a fallback or for genuinely complex tasks — but the decision of what goes where is yours, built into the routing layer, not an afterthought.

This is particularly important for NZ organisations operating under the Privacy Act 2020. Architecture that keeps data local isn't just good practice — for many use cases, it's the only compliant option.

  • Local inference — data doesn't leave your network
  • Routing logic you control, not the vendor
  • No dependency on any single cloud provider
  • NZ Privacy Act 2020 compliant by architecture
  • Full audit trail of what went where
  • Costs are predictable — no per-token billing surprises
  • Available 24/7 regardless of provider outages

Ready to build infrastructure
that actually holds up?

We'll scope your requirements and design an architecture that fits your workload, your budget, and your sovereignty requirements.

Get in touch