← blog · September 17, 2026

Trust the output, not the model: a last line of defense for automated content

When we built an automated content pipeline, the hard part was never writing the text but making sure it could never leak internal details. Here is why telling the model not to leak is not enough, and why a deterministic output filter, a fail-closed auth check, and a human approval step are the layers that actually hold.

For a while now, part of our blog has been produced by an automated pipeline. A generator drafts the text, sends it to an endpoint, and the endpoint drops the draft into an approval queue. It sounds simple. But when we built this, the hardest part was never writing the text. It was making sure the text could never leak the things it must never leak, no matter what it contained.

The first instinct is to add a line to the generator's instructions: do not write internal addresses, passwords, or server names. We tried that. The problem is that an instruction is not a security control. An instruction is probabilistic. It holds today and gets forgotten tomorrow in the middle of a long context. On top of that, a generator can carry content from its sources without realizing it. Ours was fed from change logs and technical notes, and those notes contained network addresses, container names, and internal paths. Telling the model not to leak did not guarantee it would not leak.

So we inverted the rule: trust the output, not the model. Every piece of generated text passes through a deterministic filter before it gets anywhere near publication. The filter does not deal in probability, it matches patterns. If a pattern hits, the text is rejected with a reason.

What we look for settled over time. Private network ranges came first: the 10 block, the 172.16 to 172.31 range, and the 192.168 range, plus the carrier grade CGNAT range that hands out internal addresses on mesh networks. Then came secret patterns. A private key header (a "BEGIN PRIVATE KEY" line) is a clean signature. So are access tokens: the ghp and github_pat prefixes from GitHub, the glpat prefix from GitLab, the Slack tokens that start with xox. Those prefixes are public knowledge, which is exactly why a regular expression catches them so easily. Finally we blocked internal service names. The blog's own domain is allowed, every other subdomain is not, so even if a note about which tool runs where slips into a draft, it cannot go out.

While adding this filter we kept reminding ourselves of one thing: this is the last line of defense, not the only one. We still tell the generator not to leak, and we still try to clean the source notes. The filter is the net that catches what those two layers miss. Stacking layers means one covers what another drops.

The second lesson was about authentication, and it was a classic trap. The endpoint is protected by a bearer token. In the first version, if the token was not configured, meaning it was an empty string, an incoming empty token matched the expected empty token. Forgetting to configure the endpoint meant opening it to everyone. We fixed it with a rule: if the token is not defined, the endpoint is closed. In security the default must always be closed. Leaving something open must be a deliberate decision, not the result of forgetting.

The third layer was a human. Nothing that passes the endpoint publishes itself. Every post is saved in a pending state. We love the speed of automation, but rolling back something automation published costs far more than never publishing it in the first place. A person saying yes is a cheap insurance policy.

We also closed a silent overwrite trap. If the generator sends a post twice under the same address, and that address already belongs to a published post, we do not want to overwrite the old one and drop it back into the queue, because that quietly takes a live page offline. So a submission that collides with a published address is refused and reported.

The general principle from all of this: no matter how good the generator is, treat its output as untrusted. Put a deterministic check on the last step that carries text, a file, or a reply into production. Make that check fail closed, so it rejects when in doubt. Put an approval between the producer and the publisher. And most importantly, write down why each check exists, because a rule with no recorded reason is the first thing deleted under pressure.

Since we built this pipeline, the filter has done its job a few times. The generator carried an internal address while summarizing a technical note, the filter rejected it, and we cleaned the source. Nobody published anything, nobody had to roll anything back. That is the best part of a good security control: when it works, nothing happens.