Gremlin Foresight AI finds system weaknesses and verifies the fixes
Gremlin has announced the general availability of Gremlin Foresight AI. Following a successful beta, Gremlin Foresight AI has proven it can identify and address reliability risks before they become incidents, helping engineering teams move at AI speed without compromising the resilience of mission-critical production systems.

“AI-driven development means shipping code at 10X velocity; it also means 10x the opportunity for bugs, risks, and failures,” said Kolton Andrus, CEO of Gremlin.
“While the rise of AI SRE tools are great for helping teams respond to incidents faster, it’s still cleanup after something breaks. Gremlin Foresight AI finds risks proactively, delivers the fix, and verifies the fix worked by rerunning the test. This provides teams with the assurances needed to move confidently.”
Gremlin Foresight AI is built on top of the company’s proprietary Failure Atlas, which incorporates more than a decade of cause-and-effect data around how online systems fail. This specialized foundation means every recommendation is grounded in real-world failure patterns and operational experience, not generic best practices. Every fix is validated against the test that surfaced the risk, closing the loop between identifying a weakness and proving it’s resolved.
Key features of Gremlin Foresight AI include:
- Proactive risk detection: Identify potential weaknesses and failure conditions before they become incidents.
- Guided remediation: Recommend and deliver the specific fix as a configuration patch or infrastructure-as-code change.
- Continuous validation: Re-run the originating test to confirm each fix performs as intended, and repeat testing as systems change.
- Measurable resilience: Track reliability scores across services and teams to quantify improvement and prioritize further investment.
Before founding Gremlin, Kolton Andrus was responsible for the uptime of Amazon’s retail site. He then joined Netflix to build the company’s second generation of fault-injection tooling, following the popularity of their open source project Chaos Monkey. Gremlin has since pushed the discipline of Chaos Engineering beyond one-off fault-injection experiments into a comprehensive approach to world-class reliability management.
The Gremlin platform enables teams to run planned experiments, control the blast radius, and use continuous automated testing to verify that fixes remain effective over time. Reliability scores make that progress measurable across the whole organization, giving every team a shared standard to work toward, as well as giving leaders the visibility to direct investment and hold teams accountable.
“Gremlin has always tried to help engineering teams improve the resiliency of their applications by running experiments proactively to identify weaknesses in their distributed systems, but many teams didn’t have the time or expertise to do that consistently,” said Mike Dauber, General Partner at Amplify Partners. “Foresight AI is like a trainer who does the reps for you. Your systems get stronger without your team doing all the manual work.”
Gremlin Foresight AI extends their proactive philosophy into today’s AI-driven world, delivering agentic resilience by automatically identifying reliability risks earlier and automating fixes where appropriate to avoid outages, degraded services, and negative customer impact.