Writing
Research and notes on the engineering behind operational threat intelligence.
A chatbot that invents a fact says it to one person, once. A connector that invents a fact writes it into the graph, where it looks exactly like the true ones. So I let the model read, and took away its ability to decide.
18,479 attacks that beat a language model, all of them prompt injection by construction. That makes the corpus a rare thing: a technique label you already know the answer to. Deriving it from the prompt text recovers 11.4% of it.
Thirty-two bytes is enough to fingerprint a prompt attack so that reworded versions still match, and the fingerprint can be shared without sending the prompt. This is what each size on that curve buys, measured, and where the smallest option is the right call.
Prompt attacks are a new kind of threat. The tooling growing up around them is a ladder security has already climbed once.
A threat-intel assistant that says “I don’t know” is useful. One that invents a plausible answer is a liability. So I spent a day trying to make mine invent things, and wrote down what happened.
One hacking group can carry half a dozen names. Merging them is easy until you merge two different groups by accident — which I did, ten times, including two pairs of state bodies from different countries.
A knowledge graph holds its best facts in the connections between records. Vector search can only read what is written down. So I wrote the connections down.
Keeping a knowledge base current in minutes instead of half a day, by never rebuilding what has not changed.