30 Days On Call: What Grepr AI Caught Next


This is part two of our series about Grepr AI, our proactive anomaly detection capability.
A few weeks ago we wrote about what happened when we pointed the Grepr Agent at our own cluster. In its first week it caught a leaked license key, crashlooping CI runners, and a string of code-level bugs, all without us having to write a single alert rule.
One week is a demo. So we kept it running for a month, across both our development and production environments, around the clock. Here's what it found, and what it cost.
Anything the agent catches in dev is a chance to fix it before a customer ever sees it. The following are a few that it found:
The following are some examples of what the agent found and helped us fix in our production environment:
17K investigations only produced about 20 posts, which means we didn’t have to spend a lot of time investigating false positives or deal with noise, which is a huge win. A typical hour in production looks like this:
Grepr AI investigated every signal, and it closed every one while citing evidence. When the agent did post, people read it.
One night a noisy upstream query began logging every failing row as an error. Each row looked like a new pattern, so the agent investigated them one by one, tens of thousands of times. It had already worked out that they were all the same issue: it folded them into a single memory entry and posted threaded updates, not new alerts. What it couldn't do was act on that. We hadn't given it permission to update the pipeline.
With that permission, the agent could have recognized the repeating pattern and dialed it down at the source, instead of spending time and tokens investigating every repeat. That's the enhancement we're adding now: letting the agent tune the pipeline that feeds it.
For context, Grepr AI uses 3rd party LLMs for running inference on the interesting signals that Grepr’s Autonomous Telemetry Pipelines identifies, allowing Grepr AI to run these investigations at scale. In our case, we’re using Gemini as our LLM, and running the agent around the clock across both our development and production environments initially cost us ~$100 a day. As we continue tuning Grepr AI, the cost is falling: last week came in ~$90. And now that we’re giving Grepr AI the ability to tune its sources (mentioned higher up in this blog), we anticipate costs continuing to fall, making Grepr AI an even more obvious choice.
We’re actively seeking design partners to provide feedback on what we’re building. In exchange, you’ll get to run Grepr AI for free. If you’d like to join the program, reach out to us.
What is Grepr AI?
Grepr AI is a semantic anomaly detection capability within Grepr's Autonomous Operations Platform. It reads the signals surfaced by Grepr's Autonomous Telemetry Pipeline, performs semantic analysis on the interesting ones instead of matching static rules, and decides what to do next, with access to tools over MCP and integrations like Slack and PagerDuty for escalation.
How is Grepr AI different from other AI SRE agents?
Most AI SRE agents start from a human-built alert, so they only catch failures someone already predicted. Grepr AI is triggered by Grepr’s telemetry pipeline watching every signal in the stream, so it catches unknown unknowns, the problems you didn’t write an alert for.
Do I have to write alert rules or tune thresholds to use it?
No, Grepr AI identifies interesting signals in your telemetry pipeline autonomously: when the pipeline sees a log pattern it has not seen before, that wakes Grepr AI to investigate.
Does Grepr AI escalate everything?
No, and that’s by design. Grepr AI only alerts you when something requires further attention. When an investigation determines that no further action is required, Grepr AI closes out the investigation and attaches its reasoning, so you can check its work rather than trust a black box.