How a Global Retailer Stopped Paying to Index Noise

75%
reduction in total cost of ownership
78%
fewer bytes sent to New Relic
300K+
log messages per second handled at peak
"Grepr helped us automate what to keep and what to skip, so we’re not paying to store or index noise. It lets us find the needle in the haystack without paying for the haystack!"
Evan Robinson, CTO atJitsu

How Jitsu Cut Logging Costs by 90% While Managing Millions of Shipments Generating 400 Logs Each

"Since we deployed Grepr, we’re seeing a 95% reduction in log volume and didn’t have to change a thing in our app. I'd recommend Grepr to any team that's experiencing rising costs from an expensive logging platform!"
Dave Bortz, VP Engineering at FOSSA

Case Study: How FOSSA Reduced Their Logs by 95% Without Burdening Their Engineers

“Engineers didn't change how they work at all. Dashboards and alerts still worked as expected. We just stopped paying for 90% of our log volume that was never doing anything for us.”
Ben Ede, Director of Engineering at Envoy

9 Days from Kickoff to Production: How Envoy Cut Log Volume by 90%

“After seeing what Grepr did to reduce our logs noise, extending it to traces was an easy call. We were seeing the same pattern: a lot of volume, most of it not particularly useful, and a Datadog billing model that scaled with every new host we spun up.”
Dave Bortz, VP Engineering at FOSSA
Learn More
“Engineers didn't change how they work at all. Dashboards and alerts still worked as expected. We just stopped paying for 90% of our log volume that was never doing anything for us.”
Ben Ede, Director of Engineering at Envoy
Learn More
"Grepr helped us automate what to keep and what to skip, so we’re not paying to store or index noise. It lets us find the needle in the haystack without paying for the haystack!"
Evan Robinson, CTO atJitsu
Learn More
"Since we deployed Grepr, we’re seeing a 95% reduction in log volume and didn’t have to change a thing in our app. I'd recommend Grepr to any team that's experiencing rising costs from an expensive logging platform!"
Dave Bortz, VP Engineering at FOSSA
Learn More
Company
Global Retailer
Industry
Retail
Use Case
Observability cost reduction (New Relic log volume)

The Senior Engineering Manager at a global retailer had a problem that was only going to get worse.

He oversees observability for one of the largest retail operations in the world. The platform generates between 100 and 150 terabytes of logs every month, with peaks that can exceed 300,000 messages per second. New Relic and Logz.io were doing their jobs. The bills were doing theirs too.

"Observability had become a tax," he told us. "The volume keeps growing, and the costs grow with it, but the signal we actually need hasn't changed. We were paying to index noise."

Leadership had made the mandate clear: reduce overall IT spend without sacrificing the forensic capabilities the engineering team depends on for troubleshooting and compliance. That ruled out sampling. It ruled out switching tools. And it ruled out asking developers to change how they worked.

It also, as he quickly discovered, ruled out Cribl. When the team evaluated it for their environment, it couldn't handle the retailer's peak load. At 300,000 messages per second, they needed something that wouldn't buckle.

Finding the Signal

He found Grepr and decided to test it against the real thing: production traffic, full volume, no artificial constraints.

Grepr sits between the retailer's log forwarders and their observability vendors, processing the stream in real time, identifying patterns, and routing high-signal data through to New Relic while aggregating redundant, low-value messages. Because it operates at the pipeline layer, the central team controlled the entire rollout.

The piece that made this possible without disrupting existing workflows was automatic exception handling. Grepr scanned the retailer's existing New Relic dashboards and alerts and added them as pipeline exceptions, ensuring anything currently driving an alert or populating a dashboard passed through untouched. Every dashboard kept working. Every alert kept firing. From where the development teams sat, nothing had changed.

The Constraint Nobody Talks About

There is a version of this problem that is purely technical: too much data, too high a bill, find a way to send less. But at an organization this size, the harder constraint is organizational.

Development teams have their own priorities. They are shipping features, meeting sprint commitments, and keeping their services running. A mandate from the infrastructure team to change how applications are instrumented or how logs are written lands at the bottom of a backlog nobody is motivated to clear. Even well-intentioned cost-reduction projects stall when they depend on engineering teams to take action they have no reason to prioritize.

The central observability team here understood that going in. Any solution that required buy-in or action from individual development teams was going to fail. Grepr helped this retailer avoid this constraint entirely, by allowing the central team to own and operate entirely on their own without filing a single ticket with engineering.

What Happened to the Data They Didn't Send

The other concern the team needed to resolve before committing was forensic coverage. When an incident happens, engineers need the full picture. Sampling creates gaps that surface at exactly the wrong moment, and in a retail environment with compliance obligations, incomplete log coverage is not an acceptable trade.

Grepr's answer is the data lake. All raw logs route to the retailer's own S3 bucket using open formats they fully control. The cost of storing logs there is a fraction of what New Relic charges to index them. When an incident requires a deeper look, engineers can backfill specific raw logs from the lake into New Relic on demand, pulling exactly what they need without paying to keep everything indexed all the time.

In practice, engineers rarely need to reach for it. The signal that matters is already in New Relic.

The Numbers

The results from production speak for themselves.

  • 78% reduction in bytes sent to New Relic, which directly drives billing
  • 89% reduction in message volume over sustained periods
  • 75% reduction in total cost of ownership, accounting for both vendor fees and infrastructure

The retailer has been running Grepr in production for months. In that time, the central team has received no complaints from development teams -- not because engineers were asked to accept a trade-off, but because from their perspective, nothing changed.

The Senior Engineering Manager noted, "I've been through a lot of vendor evaluations, and this was the smoothest POC we've experienced. No rules to write, no dashboards to rebuild, no conversations with developers about changing how they log. We tested it against real production traffic and it just worked."

What It Takes to Handle This Scale

Most observability cost reduction tools are built for teams running millions of log messages a day. This retailer runs that through their system in seconds. The fact that Grepr absorbed that load without configuration changes or performance degradation was, in the Senior Engineering Manager's view, the proof point that made the decision easy.

The observability bill is no longer growing faster than the business. The data is still there when they need it. And the central team didn't have to ask a single developer to change anything.

Build an engineering team that never loses knowledge.

Ready to reduce your observability TCO by 75%?

Reduce telemetry noise in your observability tools. Instantly search or backfill raw data.