Prompt injection at the tool boundary
A test harness for agentic assistants that treats every tool call as untrusted input. Surfaced a class of indirect injections that survived retrieval filtering.
I work at the seam between machine learning and security — red-teaming models, hardening the pipelines they run on, and turning findings into defenses teams can actually ship.
A test harness for agentic assistants that treats every tool call as untrusted input. Surfaced a class of indirect injections that survived retrieval filtering.
Instrumented inference gateways to flag distillation-shaped query patterns, cutting time-to-detection on scraping campaigns from weeks to hours.
End-to-end provenance for training artifacts: signed checkpoints, attested build steps, and a policy gate that refuses unverified weights at deploy.
Measured how few corrupted examples it takes to install a durable backdoor in a domain-tuned model, and which data audits actually catch it.
Blake found the failure we had already convinced ourselves was theoretical, then wrote the test that keeps it from coming back.
Rare combination: reads papers closely, and still ships tooling the on-call team is happy to own.
The clearest threat model review our AI team has had. Every finding came with a fix we could schedule.
We brought Blake in for a two-week review and kept him for the quarter. The threat model he left behind is still how we run design review.