Safety in AI isn't a checkbox — it's an architecture decision made before the first line of code. We build systems with hard constraints, verified outputs, and human escalation paths built in from the start.
The most common AI safety issues — hallucinated information, inappropriate responses, no human escalation path, failure to recognise edge cases — aren't model problems. They're design problems. They happen when safety is treated as a filter at the end, not a constraint at the beginning.
We've investigated documented cases where AI platforms flagged hundreds of high-risk user messages and took no intervention action — optimising for engagement while users were in crisis. That's not a model limitation. That's a product decision. The technical capability to do better exists. Guardian is the proof.
Built in response to documented AI safety failures. Deployed in production. The numbers speak for themselves.
NZ-Specific Crisis Detection Classifier — Firebird Solutions, September 2024
Guardian was built because the alternative was unacceptable. Existing AI platforms were demonstrably failing at crisis detection — not because the technology was incapable, but because it hadn't been built with that specific purpose in mind.
The system fine-tunes a local Qwen2.5-7B model on a custom synthetic dataset built around New Zealand-specific mental health patterns, crisis language, and normalised abuse terminology that general-purpose models miss entirely. It recognises suicidal ideation patterns, escalates appropriately for human review, and provides verified NZ crisis resources — with a hard constraint that prevents any fabricated contact information under any input.
The zero-hallucination constraint on crisis resources isn't a prompt. It's an architectural guarantee built into how the model was trained and verified. When someone is in crisis, "I made up that phone number" isn't an acceptable failure mode.
Guardian demonstrates that effective, production-ready crisis detection is technically feasible, financially viable, and buildable in weeks — not years. The argument that it's too hard has been answered.
Certain categories of failure are unacceptable regardless of input. We build these as architectural constraints — trained into the model, enforced at the output layer, verified during evaluation. Not prompt engineering. Not a filter. A guarantee.
Every AI system that touches high-stakes decisions needs clear paths for human review. We design escalation logic that's triggered reliably, routes to the right people, and maintains an audit trail. Edge cases that the model isn't confident about reach humans — not a 500 error.
Verification systems that check model outputs against defined criteria before they reach the user — factual constraints, format requirements, safety guardrails. Particularly critical for any output that references real-world resources, contacts, or factual claims.
Evaluation of existing AI deployments against safety criteria — where are the failure modes, what inputs produce unreliable outputs, what's the escalation path when it goes wrong. Honest assessment, not validation.
Two articles documenting AI safety failures and the case for doing better. Published on DEV.to.
An investigation into AI crisis response failures and the corporate decisions that allowed them to persist. The documented cases, the technical analysis, and the argument that this was preventable.
DEV.to — Technical WritingOn AI companies that prioritised commerce over safety infrastructure — and what the prioritisation of engagement metrics over user protection looks like from the inside of the codebase.
We design AI safety in from the start. Tell us what you're building and what the failure modes are — we'll design an architecture that handles them.
Get in touch