The archive · AI & Models · Product decision · 2026
Silmaril bets prompt-injection defense must self-heal: YC W26 firewall blocks in 20ms
YC Spring 2026 startup Silmaril sells a runtime firewall for AI agents that blocks harmful tool calls in ~20ms and retrains within an hour of each attack.
Silmaril
What the business is
Silmaril sells runtime security for AI agents: a firewall that evaluates intent, application context, and execution state, blocks harmful tool calls in about 20 milliseconds, and retrains itself within an hour.
How it started
Aum Upadhyay built security frameworks protecting AWS's AI infrastructure, and Eduardo Velasco developed hardened models for Amazon's homepage; after leaving, the pair became whitehat hackers and, by their account, hacked OpenAI, Anthropic, Google, and Microsoft 15 times in two weeks. Seeing how far ahead attackers were, they founded Silmaril in San Francisco in 2026 in Y Combinator's Spring 2026 batch.
What happened
Silmaril launched in 2026 claiming it blocked 96% of real-world contextual and emerging attacks versus 61% for leading guardrails, with about 20ms latency overhead and a self-healing loop that retrains and deploys updated weights within an hour. YesPress reported 95.6% accuracy across 131 attack techniques versus 79.2% for Perplexity BrowseSafe and 74.3% for Lakera GuardAI, and the company said it had stopped $28M of customer damages.
No ending yet — it is still running.
Background
Silmaril, a San Francisco startup in Y Combinator's Spring 2026 batch, bets that prompt-injection defense cannot be a static filter. As AI agents read web pages, emails, and documents and then act on them — calling tools, moving money, touching databases — a sentence slipped into an untrusted document can borrow the agent's hands, so the company asks a different question than the guardrails: not 'is this input bad?' but 'is this execution heading somewhere harmful?'
Co-founders Aum Upadhyay and Eduardo Velasco came from opposite sides of the same discipline: Upadhyay built security frameworks protecting AWS's AI infrastructure, and Velasco developed hardened models for Amazon's homepage. After leaving, they became whitehat hackers and, by their account, broke into OpenAI, Anthropic, Google, and Microsoft 15 times in two weeks, which convinced them attackers were ahead of every existing defense.
The product wraps agents in a runtime firewall that reads user intent, application context, and execution state, blocks a harmful tool call in about 20 milliseconds, and runs autonomous threat-hunting agents that probe each customer's app for multi-step attack chains. Verified exploits feed a retraining loop that updates deployed weights within an hour, with anonymized learning shared across deployments.
The company claims 96% block rates against real-world contextual threats versus 61% for leading guardrails and $28M of customer damages stopped, though the figures are self-reported. Silmaril is early — a team of roughly two to three — and competes with Lakera, Prompt Security, Protect AI, and HiddenLayer in a crowded AI-security gold rush.
What has to be true
- Prompt injection is the top unsolved risk for tool-calling agents, a market that barely existed before agents shipped.
- Outcome-based detection catches multi-step attacks that single-input classifiers miss by design.
- 20ms latency makes the control invisible enough to stay enabled — the difference between a control that ships and one that gets disabled.
- The self-retraining loop plus anonymized threat sharing compounds in value with every deployment, a network effect security products rarely get.
What can be applied
In a brand-new failure mode, the defensible claim is speed plus learning: a control that adds 20ms and retrains itself fits deployment reality, while filters that need manual tuning get switched off.
Aftermath
As of September 2026 Silmaril is seed-stage with a team of roughly two to three, integration through language SDKs, agent plugins for Claude Code, OpenClaw, and OpenCode, and AI gateways like LiteLLM, sold as SaaS or self-hosted. It cites AICPA SOC and ISO 27001 certifications and memberships in the Cloud Security Alliance and NVIDIA's Inception program, but performance and damage figures remain the company's own, with independent benchmarks still pending.
Sources
- Launch YC: Silmaril: Self-healing Prompt Injection Defense
- Silmaril: The Self-Healing Firewall for AI Agents
- Silmaril: Security for agents that self-improves
spotted an error? The archive wants to know.
Your turn
You just read one. Describe what you are building, and see who is betting on the same thing.
Free account · 3 free questions · no card