Back
Author
Utkarsh Vashishtha
LAST UPDATED
October 5, 2026
Description
In its first month running 24/7 across our development and production environments, Grepr AI ran about 17,000 investigations, pinged us in Slack only when it mattered, and caught leaked secrets, crash loops, and code-level bugs, all without a single alert rule.
Engineering Guides

30 Days On Call: What Grepr AI Caught Next

Blog post featured image
IN THIS ARTICLE
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..

This is part two of our series about Grepr AI, our proactive anomaly detection capability.

A few weeks ago we wrote about what happened when we pointed the Grepr Agent at our own cluster. In its first week it caught a leaked license key, crashlooping CI runners, and a string of code-level bugs, all without us having to write a single alert rule.

One week is a demo. So we kept it running for a month, across both our development and production environments, around the clock. Here's what it found, and what it cost.

The month in numbers

  • ~17K investigations in production (more in dev), each triggered by a log pattern the pipeline hadn't seen before.
  • ~20 Slack posts from Grepr AI, alerting us to issues that actually needed our attention (the rest were non-issues, where Grepr AI closed each one with its reasoning attached).
  • ~4–5 minutes, typically, from the first signal to a diagnosis/fix in Slack.
  • Zero alert rules written.

‍

Caught in dev, before it shipped

Anything the agent catches in dev is a chance to fix it before a customer ever sees it. The following are a few that it found:

  • A broker migration, de-risked. We were trialing a new message broker in dev. The agent caught TLS handshake failures between our services and the new brokers, and later a controller crash-looping on restart. That second one led an engineer to a misconfigured setting. Both surfaced in dev, where they cost us nothing.
  • Auth codes leaking into access logs, fixed by the agent. It noticed one-time sign-in codes showing up in our web server's access logs, traced them to the browser's Referer header, and opened the fix PR itself. Then it checked the running build and flagged that the fix still needed a redeploy.
  • An auth token in an exporter's error logs. While investigating something else, it noticed one of our telemetry exporters prints its ingest token in the URL whenever a request fails.
  • A query bug, found at the same time as a human. It traced a failing search to the exact method that couldn't handle it. The engineer who owned that code replied: "Good bot… I just found this and put up a PR for it."

‍

Caught in production, in minutes

The following are some examples of what the agent found and helped us fix in our production environment:

  • Intermittent ingestion errors. Few requests started failing with memory saturated errors. Within five minutes the agent had named the blocking call, ruled out a deploy or a code change, and pointed to the message broker the requests were waiting on. Although the failures were retryable and transient, it helped us improve the QoS by helping us right-size our infra.
  • Memory for a recurring problem. When an ingestion pod ran out of worker threads, the agent's memory linked the event to a dozen earlier occurrences on other pods. The service owner replied, "Did we never end up figuring out a solution to this?" It had an owner that day.
  • Secrets in logs. It caught a setup job printing a broker password to its logs, and a code path logging invitation tokens in plaintext. It opened a fix PR for each.
  • A release that referenced a missing image. New pipeline deployments started failing with ErrImagePull. In about 15 minutes the agent traced them to an image that had never been pushed to the registry.
  • A crash loop that changed shape. A rollout failed first on half-created resources, then hours later on a secret scheduled for deletion. The agent connected the two, and found a latent exception-handling bug along the way.

‍

Knowing when to say nothing

17K investigations only produced about 20 posts, which means we didn’t have to spend a lot of time investigating false positives or deal with noise, which is a huge win. A typical hour in production looks like this:

  • ALL CLEAR: 1 signal, noise. Routine autoscaling audit log at INFO.
  • ALL CLEAR: 2 signals, both noise. CI runner cleanup at INFO.
  • ALL CLEAR: 1 signal, noise. Handled validation error returning HTTP 400.

Grepr AI investigated every signal, and it closed every one while citing evidence. When the agent did post, people read it.

‍

What we're improving: Allowing agents to tune the pipeline

One night a noisy upstream query began logging every failing row as an error. Each row looked like a new pattern, so the agent investigated them one by one, tens of thousands of times. It had already worked out that they were all the same issue: it folded them into a single memory entry and posted threaded updates, not new alerts. What it couldn't do was act on that. We hadn't given it permission to update the pipeline.

With that permission, the agent could have recognized the repeating pattern and dialed it down at the source, instead of spending time and tokens investigating every repeat. That's the enhancement we're adding now: letting the agent tune the pipeline that feeds it.

‍

The falling LLM costs

For context, Grepr AI uses 3rd party LLMs for running inference on the interesting signals that Grepr’s Autonomous Telemetry Pipelines identifies, allowing Grepr AI to run these investigations at scale. In our case, we’re using Gemini as our LLM, and running the agent around the clock across both our development and production environments initially cost us ~$100 a day. As we continue tuning Grepr AI, the cost is falling: last week came in ~$90. And now that we’re giving Grepr AI the ability to tune its sources (mentioned higher up in this blog), we anticipate costs continuing to fall, making Grepr AI an even more obvious choice.

‍

What we learned (part two)

  • Discover in Dev, avoid issues in Prod. Run the agent wherever your code lands first, not just in production, so that you can proactively remove bugs before they impact customers.
  • Memory turns alerts into history. "This has happened 12 times before" is worth more than any single alert.
  • An agent should be able to tune its own inputs. An agent that can only investigate a noisy pattern pays for that pattern over and over. One that can dial it down pays once.

‍

Try it out: Design partnership is open

We’re actively seeking design partners to provide feedback on what we’re building. In exchange, you’ll get to run Grepr AI for free. If you’d like to join the program, reach out to us.

‍

FAQ

‍What is Grepr AI?
Grepr AI is a semantic anomaly detection capability within Grepr's Autonomous Operations Platform. It reads the signals surfaced by Grepr's Autonomous Telemetry Pipeline, performs semantic analysis on the interesting ones instead of matching static rules, and decides what to do next, with access to tools over MCP and integrations like Slack and PagerDuty for escalation.

How is Grepr AI different from other AI SRE agents?
Most AI SRE agents start from a human-built alert, so they only catch failures someone already predicted. Grepr AI is triggered by Grepr’s telemetry pipeline watching every signal in the stream, so it catches unknown unknowns, the problems you didn’t write an alert for.

Do I have to write alert rules or tune thresholds to use it?
No, Grepr AI identifies interesting signals in your telemetry pipeline autonomously: when the pipeline sees a log pattern it has not seen before, that wakes Grepr AI to investigate.

Does Grepr AI escalate everything?
No, and that’s by design. Grepr AI only alerts you when something requires further attention. When an investigation determines that no further action is required, Grepr AI closes out the investigation and attaches its reasoning, so you can check its work rather than trust a black box.

‍

Ready to reduce your observability TCO by 75%?
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..