← build journal
Entry 1011 January 2026

Long-Context Model Security

My notes on many-shot jailbreaking and why long context changes the security boundary around model-assisted workflows.

LLM SecurityJailbreakingSecurity Research

Research notes

I spent this note looking at many-shot jailbreaking. Longer context windows expand the attack surface as well as capability, and repeated examples can steer a model away from intended behaviour even when individual prompts appear harmless.

This matters for enterprise systems that combine user input, retrieved documents, and long conversation history. Security controls must consider the full assembled context rather than validating only the latest message.

Product implication

Language models used in threat analysis should not control evidence, attribution, or release decisions. They can help organise information, but deterministic checks and cited sources need to remain authoritative.

The question I care about is whether the surrounding system can bound the model’s influence and fail safely when inputs are hostile or confidence is low.