Guardrails before autonomy. Always before.

Series: The Build Track · Post 6 of 7

The chapter with our incident in it: a local model fabricated client case studies and two posts auto-published. The root cause, and the five layers that exist now.

One morning my system published two posts containing client case studies that do not exist.

One went to my LinkedIn business page, one to my Facebook page. A local model had generated them. Fake nonprofits, fake clients, fake percentage results, written with total confidence in my voice. Thirty-one more posts like them were sitting in the queue. Five were scheduled for that same day.

The system caught it the next morning. The daily planner read the queue, recognized the fabrications, froze both pages before the day's dispatch window, and wrote up the incident with the fix. Then the alerting layer failed too. The Telegram bot was unreachable from where the planner runs, so the alert that should have paged me never sent. I found out from the morning brief.

That's the honest version. Now the useful part: why it happened, why it was an architecture failure and not a bad prompt, and the layers that exist now. If you're building one of these, this chapter is the one to steal from. If you're evaluating someone else's, this chapter is the interrogation script.

The root cause was boring, which is the lesson

The fabricating model is the surface cause. The real failure was a config file.

Our operating envelopes, the per-platform caps I'll get to in a minute, are defined per account. The accounts in that config still had placeholder ids from initial setup. So when the dispatcher checked each queued post against its envelope, it found no matching account and silently skipped enforcement. Every gate downstream of that lookup quietly waved everything through. The same queue held three LinkedIn posts for a single day against a one-per-day cap, and nothing objected.

Read that failure chain again, because it generalizes. The unsafe content didn't beat the guardrails. The guardrails were configured off and nothing complained about it. A check that silently passes when its config is broken isn't a check. Fail-open is the default behavior of code nobody thought hard about, and fail-open is how two fabricated posts reach a real business page.

There was one more compounding finding, and it stung. While the local model was generating slop, 45 hand-written, voice-scored posts sat in a content library completely unused. The pipeline had been pointed at the wrong source the whole time. The immediate fix was one command that cancelled the generated queue, ingested real library content, and lifted the freezes, plus a standing morning step to cancel anything the generator queued overnight until the generation path itself was fixed or disabled. Interim guardrails are allowed to be a human running a script. They are not allowed to be hope.

So: the layers. Four of them, plus the architectural rule that makes them enforceable.

Layer 1: approval queues. Trust is earned per channel

Tier A policy, written in the envelope spec: 100% of public-posting actions go through an approval queue. A human sees it before the world does.

I think of approval queues as the autonomy throttle, not a permanent state. The progression is full review while a channel proves itself, then sampling, then spot-checks. Any incident resets the channel to full review. The mistake people make is binary thinking: either a human approves everything forever (so why build the machine), or the machine is autonomous on day one (see: the incident above). The queue is how you move between those states deliberately instead of by accident.

The honest cost: queues back up. Ours hit roughly 163 pending items during a stretch when I wasn't reviewing daily, which the system flagged, correctly, as the worst of both worlds. Generation on, review off. If you build a queue, you're committing to working it, or to explicitly lowering generation until you can. The queue's length is itself a metric the system should be watching.

Layer 2: operating envelopes. Caps with two ceilings

Every action type on every platform has an envelope with two numbers: a soft target the system hovers near on a normal day, and a soft max it never exceeds without an explicit waiver. Behind both sits a documented reference to the platform's actual ceiling, so the numbers trace to something.

Concretely, from our config. LinkedIn company posts: soft target 1 per day, soft max 1, because LinkedIn penalizes more than one post per 18 hours. Facebook page posts: target 1, max 3, with a minimum gap between posts. X: target 3, max 8. Around the counts sit shape constraints that keep behavior plausible rather than mechanical. Active hours 7am to 10pm, a silent overnight window, randomized gaps between actions, no duplicate text within 30 days.

The envelope spec also defines automatic freeze conditions. Two consecutive rate-limit responses, or any block screen, freezes the account entirely. The platforms' own immune systems are treated as tripwires, not obstacles.

Two ceilings matter because they separate the tuner from the limiter. Chapter 5's training loop is allowed to move behavior around below the soft max. Nothing automated moves the max. Optimization pressure and safety limits must not live in the same variable, or the optimizer will eventually eat the limit.

Layer 3: the fabrication gate. Never trust a prompt to enforce a rule

Our generation prompt contains a rule in plain language: only use numbers, client names, or results that appear verbatim in the provided source material; never invent a client or an outcome; if you have no facts, write opinion content with zero specific claims.

I want that rule in the prompt. I do not trust it, and neither should you. A prompt is a request. The incident is what a request is worth under pressure.

So there's a mechanical gate that doesn't ask the model anything. After generation, plain regex extracts every specific claim from the output. Percentages, dollar amounts, multipliers like "3x," count claims like "200 leads," client-result sentence patterns like "our client doubled..." Each one gets checked against the exact source text the model was given. Any claim not present in the input is flagged as invented. And the code comment states the policy outright: flagged posts can never auto-queue. They land in draft, flags attached, for human review. It's an absolute override. Even a post that passes every quality, safety, and voice check stays held if it contains one unsubstantiated number.

The principle generalizes past marketing: wherever an LLM's output crosses into the world, put a deterministic checker between them. The checker doesn't need to be smart. Regex against source text is almost insultingly dumb. It would have caught both of those posts, because a fabricated "grew 40%" contains a percentage that appears nowhere in the inputs. Dumb and mechanical beats smart and probabilistic at the boundary, every time.

Layer 4: kill switches. Four scopes, one semantics

When something goes wrong, you need stops at more than one radius. Ours are four flag files on disk:

  • KS-1, global: one flag pauses every outbound action the system can take.
  • KS-2, per-account: freeze one account. This is what the planner tripped for both business pages.
  • KS-3, per-persona: freeze every account belonging to one identity.
  • KS-4, network-wide: the everything-in-a-category stop for the experimental tier.

They're checked broadest to narrowest on every single action, and the decision, which switch, which flag file, what reason, is written into the action's log row, so a skipped action is auditable later.

Why flag files instead of database rows or API calls? Because a kill switch must work when everything else is broken. A file on disk can be created over SSH, by a cron job, by me in a panic, with no daemon healthy. The semantics are deliberately simple. File present and non-empty means active. Even clearing one has a fallback: if the file can't be deleted, truncating it to zero bytes counts as cleared. The off switch has its own off switch.

And the docstrings are honest about maturity. Two of the four levels are fully wired end-to-end; the other two are checked on every action but their automatic trip rules ship with a later phase. Guardrails have version numbers too. Pretending otherwise is how you get surprised.

The one choke point, and the question that exposes everyone

None of the four layers means anything without an architectural commitment: every outbound action passes through exactly one function.

In our dispatcher, directly in front of every platform adapter's post call, sits a single gate. It checks the kill switches in order, resolves the account's persona so a persona freeze binds even when the caller didn't pass one, and returns allowed-or-blocked with a named reason. The code's own comment says it plainly: every posting path must pass through this exactly once, at the lowest dispatch level. Not in the planner, where a clever path can route around it. At the last possible moment before the network call. Budget and envelope checks live at the same choke point, and the post-incident rule is that they fail closed. A check that errors blocks the action. It never waves it through.

The placeholder-id failure happened because a check effectively evaluated "account unknown" as "no limits apply." Fail-closed inverts that: unknown account, no action. You will occasionally block something legitimate. That's the correct trade. The alternative is what the incident cost.

So here's the test, and it's the single most useful question I can give you for evaluating any agent product, framework, or contractor's build: "Show me the one function every outbound action passes through." If the answer is a function, with a file and a line number, you're talking to someone with guardrails. If the answer is a diagram, a philosophy, or "each integration handles that," you're talking to someone who hasn't had their incident yet. Everyone gets one. The architecture decides whether it's a story you tell in a blog post or a lawsuit you settle quietly.

Previous
Previous

Build, buy, or hire: the honest math

Next
Next

Training modules: improving on a schedule, not by magic