Skip to content
Kentron Technologies

AIEngineeringReliability

Guardrails: how to stop an AI feature embarrassing you in production

The demo always works. What decides whether an AI feature survives is the logging, the evaluation and the plain path for when the model is unavailable or wrong.

Author
Kentron Technologies
Published
Reading time
6 min read
Working at a laptop

An AI feature is easy to build and hard to keep. The demo works because the demo was tried a dozen times with inputs the builder chose. Production supplies inputs nobody imagined, at a volume that makes rare failures routine, and does it on a day when the model provider is degraded. This article is about the unglamorous machinery that decides whether an AI feature is still switched on a year later.

Three things every AI feature needs

We apply the same three rules to every model-backed feature we ship, in our own products and in client work.

  1. A log of what it did. The input, the output, the model version, the latency and the cost, for every call.
  2. A plain fallback. A defined, non-AI path for when the provider is down, slow or refuses.
  3. A human with authority. For anything consequential, the model drafts and a person approves.

None of this is exotic, and all three are routinely skipped because they are not visible in a demo. The first one is the foundation: a feature without logs cannot be debugged, improved or defended when a customer disputes what happened.

Why the model version belongs in your log

Providers update models. A prompt tuned against one version can behave differently against its successor, usually subtly: slightly longer outputs, a changed willingness to refuse, a different format. Without the model version in your logs, that shift is almost impossible to diagnose, because the symptom appears on a day you changed nothing.

Where a provider offers pinned versions, pin them, and treat a version upgrade as a change requiring the same testing as a code change. Behaviour that arrives without a deployment is the hardest kind to explain.

Evaluation, at the smallest scale that is honest

Formal evaluation frameworks are worth it at scale. For most projects the minimum viable version is a file of cases and expected behaviours, run before each release.

  • Collect twenty to fifty real inputs, including the ones that went wrong. Production failures are the most valuable test cases you will ever get.
  • Record what good output looks like for each, in a sentence if it cannot be asserted exactly.
  • Include cases where the correct answer is a refusal, so you can detect when the model stops declining.
  • Re-run them after any prompt change, model change or retrieval change.

Fifty cases in a file, reviewed by a human in twenty minutes, catch most regressions. A team that tells you evaluation is manual has at least thought about it; a team with no answer has not.

Your best test cases are your past failures. Keep them.

Designing the fallback

The fallback should be the thing the user would have done anyway, reachable without them understanding that a model failed.

FeatureFallback
Voice agent on a phone lineRoute to a human queue, or take a message
Dictation to draft a summaryThe ordinary typing form, pre-filled with what exists
Document extractionManual entry on the same review screen
Chat answering questionsThe contact number, offered immediately

Two rules make this work. Set a timeout and honour it, because a model call that takes thirty seconds has already failed from the user's point of view. And never show a raw provider error; a front desk cannot act on a rate-limit message.

Guarding what the model is allowed to do

For agents, which take actions rather than only producing text, the guardrail is the permission boundary. The set of tools an agent may invoke should be declared explicitly and enforced server-side, never inferred from the prompt. When we built our voice agent platform, each tenant declares the tools its agent may call and the platform executes them against that tenant's own endpoints, with tenant scoping applied at the query rather than in the interface.

Two further limits are worth building in from the start: a cap on how many actions one conversation may take, which contains a loop, and confirmation before anything irreversible. An agent that can cancel an order should say what it is about to cancel first.

Watch cost as a signal, not just a bill

Per-conversation cost is a diagnostic. A sudden rise usually means something is wrong: retrieval returning too much context, a loop, retries, or an abusive pattern. Alert on cost per conversation rather than only on the monthly total, and you will find bugs before finance does.

A short pre-launch checklist

  1. Every call logged with input, output, model version, latency and cost.
  2. A timeout, and a fallback a user can act on without knowing why.
  3. A case file of real inputs, including failures and expected refusals.
  4. Explicit server-side limits on what an agent may do, plus a cap per conversation.
  5. Confirmation before anything irreversible.
  6. An alert on cost per conversation and on error rate.
  7. A documented way to switch the AI path off without a deployment.

That last item has saved more reputations than any prompt. A feature flag that disables an AI path in seconds turns an incident into an inconvenience.

Frequently asked questions

What are AI guardrails?

The constraints and machinery around a model that keep it useful and safe in production: logging every call with its model version, a defined non-AI fallback, explicit server-side limits on what an agent may do, human approval for consequential actions, and alerts on error rate and cost. The model is the small part; the guardrails are what make it dependable.

How do I test an AI feature before release?

Keep a file of twenty to fifty real inputs with their expected behaviours, including inputs that previously went wrong and cases where the right answer is a refusal. Re-run it after any change to the prompt, the model or the retrieval layer. Past production failures are the most valuable test cases available, so collect them deliberately.

What should happen when the AI provider goes down?

The user should reach the path they would have used anyway, without needing to know a model failed: a human queue for a voice agent, an ordinary form for dictation, manual entry for extraction. Set a timeout and honour it, never display a raw provider error, and keep a feature flag that disables the AI path without a deployment.

Why does an AI feature that worked suddenly behave differently?

Usually because the provider updated the model underneath you. A prompt tuned against one version can produce different lengths, formats or refusal behaviour on its successor, and the symptom appears on a day you changed nothing. Log the model version with every call, pin versions where the provider allows it, and treat a version change like a code change.

Kentron Technologies

Editorial team

Builds and runs Kentron Technologies’s products. Writes here when a decision was hard enough to be worth explaining.

Next step

Tell us what you are running, and what is slow.

A demo of any product, or a conversation about something that does not exist yet. Either way, you will talk to someone who builds the software.

CallWhatsAppTalk to us