Writing
A follow-up to the Model Trust Gate. I put the method to a real, loaded test: a French, a Chinese, and an American open model, run through the same evaluation. Each failed something, and none failed the way its origin would predict. The case for judging a model by testing it, not by where it was made.
A methodology, not a product: a seven-layer, fail-fast gate that turns 'do we trust this model for this use?' into a repeatable, auditable decision, composing NIST, ISO/IEC 42001, CSA AICM, and OWASP rather than inventing new controls.
A post-mortem. The MCP scanner in my local red-team pipeline was defeated by a supply-chain problem, the exact class of risk security tooling exists to catch. One word, two package ecosystems, an acquisition, and weeks of quiet false confidence.
The Omnibus is set to push standalone high-risk enforcement to December 2027. Every law firm called it relief. Having built the blueprint, I read it as a deadline for a multi-quarter engineering programme — and the clock starts now.
A field report from building a fully-local, three-layer AI red-team pipeline on an Apple-silicon laptop: 28 fixes, a two-model nightly run, and the honest gap between a local 'pass' and an actual security verdict.
Notes on the NSA's May 2026 MCP Security CSI and practical defensive and adversarial work in this space.
How to spend a weekend implementing OWASP, NIST, and CSA guidance, and what I learned about where the real security boundaries live.