# How the disclosure gate on this site keeps a live trading edge silent

Canonical: <https://alwaysriskon.com/writeups/the-disclosure-gate>
Published: 2026-09-30

The disclosure gate on alwaysriskon.com is four layers deep, from the content schema to a scan of the built pages, and its keyword list reads 25 terms. It blocks the build on a banned word, and it still cannot prove that a page says nothing about a live edge, so a person ticks that box by hand.

## Key numbers

| Claim                                | Value     | Measured on                           | As of      |
| ------------------------------------ | --------- | ------------------------------------- | ---------- |
| Terms the keyword scan reads         | 25 terms  | the site's test suite, September 2026 | 2026-09-30 |
| Unit tests on the disclosure scanner | 462 tests | the site's test suite, September 2026 | 2026-09-30 |

## The game and the prize

This site publishes numbers about everything I build, and it has one rule about trading. A live trading edge gets total silence on alwaysriskon.com, and a dead one gets a full post-mortem with dates and numbers. The game is to keep that rule on every page, every feed and every file a crawler can read, while the site keeps publishing.

There is no number for the prize, because the prize is a sentence that never gets written. The cost of losing is easy to state though. An edge printed on a public page stops being an edge.

## The naive estimate

The obvious version is one keyword list and a search over the Markdown before anything goes out. The list on this site reads 25 terms, and every leak I could picture contains at least one of them, so a search over the source looked like the whole job.

**25 terms against 4 templates**

- Terms the scan reads: 25
- Templates it guards: 4

I did the math on it and it doesn't hold. A reader never sees the Markdown. They see a page built from it, plus the list pages that quote its title, the feed that carries its full text, the plain copy the site serves for AI crawlers, and the structured data in the page head. The same words reach a reader through every one of those, and a search over the source reads none of them.

## The descent

So the gate has four layers, and any one of them failing blocks the build.

The first layer is the schema. An open position cannot hold the fields a finished one uses for its post-mortem, and any key the schema doesn't list fails when the file is read. Spend cannot be a negative number.

The second layer works out what each item is linked to. A write-up takes the status of the position it links, and a write-up that links nothing is read as strict, which means it is treated as if it were open. That default matters more than it looks. The easy mistake is a new piece with no position set, and strict turns that mistake into the safe case.

The third layer is the keyword scan itself. It reads the body, the frontmatter, every YAML and chart file, the CSV headers, image alt text, captions and the text inside inlined diagrams, for every strict or open item. Key numbers are read with their value and unit joined, so a signed figure is caught as one string.

The fourth layer runs the same scan over the built site. It reads the HTML of the list pages, every position hub and every page linked to an open position, plus the feed, the AI crawler copy and the plain Markdown copy of each write-up. This is the layer that catches a word the templates put there and the source never held.

Unit tests in each disclosure test file, from vitest's own count, most of them on the scanner library. [Figure](https://alwaysriskon.com/writeups/the-disclosure-gate#fig1)

Then the letters. A word spelt with a Cyrillic o, a Greek capital, small capitals or a combining mark looks the same to you and completely different to a pattern match. Unicode's own security report catalogues these look-alikes in far more detail than any list I would have written, and hats off to them for it. The scan now strips invisible characters and folds look-alikes to plain letters before it matches anything.

Then hidden elements, which is where the gate failed. Put a banned word on a page with a hidden element in the middle of it, and the page scan read harmless pieces while the browser showed the whole word. All 15 gate steps were green when the final review found it. It was the one critical finding in that review.

The fix has a part in the content check and a part in the page scan. Raw HTML is now banned in a write-up body outright, so an author can only use the site's own components. And the page scan reads anything hidden as its own separate chunk, so hidden text is still scanned and never joined to the words around it.

The last piece is the marker contract. A mixed page, such as the list of every position, holds open and closed items side by side. Every item's outer element carries its position and its status as markers, so the rendered scan knows which rows are open and scans those. A row with no marker, or a marker that doesn't match its content, fails the build.

## Receipts

The keyword list is 25 terms, and it is the same list for every template. The disclosure code has 462 unit tests across its test files, counted by vitest on a clean checkout.

**462 tests for 25 terms**

- Unit tests on the disclosure code: 462
- Terms the scan reads: 25

462 tests for a list of 25 words sounds like overkill, until you count the ways a word can reach a page. Most of the tests are on the scanner library, the part that folds look-alike letters and strips invisible characters, because the keyword scan and the scan of the built pages both run on it, so a spelling it misses gets past both. 28 commits touched the scanner and its lint scripts before this piece was written.

## The edge, or its honest absence

A keyword scan cannot prove absence. It proves that a word is not on the page. A sentence that describes a live edge without using any of the 25 terms passes all four layers, and no pattern will ever catch it.

So the last check is a person. Every pull request carries a checkbox that says nothing on any page concerns a live edge, and I tick it myself after reading the page, never an agent. The scan catches the careless mistakes, and the checkbox is there for the sentence that says too much without a banned word in it.

*The scan proves a word is absent, never the edge.*

## What carries forward

The scanner is one library, and the content check, the build, the rendered gate and the deploy all call the same functions. Every template gets the rules unchanged, so a new template can change the structure of a page and never the rules on what it may say.

If you publish your own numbers next to something you need to keep quiet, the method copies across. State the rule in one sentence, make an unlinked piece strict by default, scan what the reader receives and not only what you wrote, fold look-alike letters before you match, and keep a person at the end.

## The cost

28 commits before this piece, and 462 tests to keep green. Text inside an image, whether a screenshot or a diagram drawn as SVG paths, is the shape none of the four layers reads, so every image and diagram on this site gets read by eye before it merges. That check is manual, and it stays manual until something better exists.

## Check it yourself

- **The terms**: 25 terms, The rule is spec section 7.1. Every term applies to every page that names an open position or none.

## Sources

1. [Unicode Technical Standard #39, Unicode Security Mechanisms](https://www.unicode.org/reports/tr39/)
