Back
Author
Summer Lambert
LAST UPDATED
September 1, 2026
Description
An engineer's three extra log lines cost Envoy $40,000 in three days. Here's how their Director of Engineering cut Datadog log volume more than 90% in nine days, without anyone noticing.
Events

Webinar Recap: Envoy cut more than 90% of their Datadog log volume in nine days

Blog post featured image
IN THIS ARTICLE
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..

An engineer at Envoy once turned on three extra lines of logging. Three days later, it had cost the company $40,000.

Ben Ede, Envoy's Director of Engineering, told that story about halfway through our webinar, almost in passing. Everyone who runs observability has a version of it. A team flips something on while they are debugging, forgets about it, and the bill keeps climbing without anyone watching.

Ben joined Envoy just under a year ago to run the infrastructure, security, and platform teams. For his first eight months the observability spend was steady. Then it doubled in about three or four months. "That was the shocking thing, how quickly we doubled," he said.

His first instinct was the obvious one: find the teams spending more than usual and go talk to them. It sort of worked. Some teams had turned on something they did not realize was expensive and switched it back off. But others needed it, just for a week, to chase a bug. "Then another cost comes in that trumps that one, the first cost gets normalized, you forget about it, and you move on to the next biggest spender. That approach doesn't scale."

The frustrating part is that Envoy already had the tool everyone points to. "We had been writing exclusion rules, that's the crazy thing," Ben said. "It just became another whack-a-mole exercise." The granularity was too fine to write a rule that killed the noise without risking the one log you would want during a P0. One service, he mentioned, was generating something like four billion logs. When he asked the engineering manager if they needed them, the answer was no.

Why it kept getting worse

Grepr's CEO, Jad Naous, framed the bigger shift before turning it over to Ben. Over the last fifteen years observability went from something you bolted onto a few services to something deep and everywhere, right as architectures broke into microservices that spin up constantly. More surface area, more telemetry, and the value of any single log dropping fast. Telemetry became insurance you pay for up front and rarely collect on.

Then AI poured gasoline on it. "Claude increased the amount of logs going live, because our productivity went up," Ben said. "But also, why use one word when you can have a whole sentence in a log?" Jad hears the same thing from other teams, usually with a twist: they want to cut the observability bill so they can spend more on tokens.

What had to be true first

Before Ben would let anything sit between his services and Datadog, it had a short list to clear. It could not hurt reliability, and it could not add latency, since some of his paths into Datadog were already slow. It had to be simple enough that his infra team would not have to babysit it, which matters more than ever now that everyone is hiring away infra people for AI ops. And after you counted the setup time, it still had to save real money. "Almost a drag-and-drop," was how he put it.

The rollout, and the *chef’s kiss* silence

"I can't remember how many minutes it took us to get Grepr working, but it was super quick," Ben said, and he took that as the first good sign.

They generated synthetic traffic and the number jumped to 80 to 85%. Envoy had already decided that anything above 70% made this worth doing. From the day-zero kickoff to full production was nine days, with Grepr Engineer, Suneet, running one-on-one sessions to get the Envoy engineering team up to speed.

The rollout is the part Ben enjoyed most. Engineers were the skeptics going in, worried that changing the logging pipeline would break their dashboards and alerts. Grepr scans your existing dashboards, alerts, and log-based metrics and adds them as exceptions, so the logs powering the things people actually look at keep flowing untouched. So when the team announced the cutover, the reaction was nothing. "There was deathly silence. No one said anything," Ben said. "A few months later there are people completely unaware we're using Grepr, because it's just doing its job."

What actually changed

Envoy landed north of 90% log reduction into Datadog, higher than Ben expected going in. But the number he did not see coming was the time. "I forgot how much time I got back," he said. No more three or four hours a week hunting for whatever engineer had just cost the company money. In his words: "So I can make more coffee." (Ben, for the record, is serious about coffee.)

One part was genuinely painful: legal. It was the slowest thing in the whole project. Envoy is a security company, mid-way through SOC 2, HIPAA, CMMC, and ISO 27001 this year, and vetting a new supplier takes time. Even though they do not log PII as a matter of practice, a customer might drop something sensitive into a field never meant for it. His advice to anyone security-conscious: get ahead of legal early. It helps that the raw data lands in Envoy's own S3 bucket, so they set their own retention policies and can delete whatever they need to. Compliance, Ben said, has not been an issue.

Ben's sage advice

Asked what he would tell someone stuck in the middle of this exact problem, Ben did not hedge. "Why wait? It's one of my biggest regrets that it took us so long to [deploy Grepr]." It took him about three months after joining to even realize Grepr could solve it.

Watch The Recap

If your log bill is growing faster than your business, you can try Grepr on the free tier, no credit card, at app.grepr.ai/signup. Most teams see their own reduction number in under thirty minutes.

Frequently Asked Questions

How much did Envoy reduce their Datadog log volume?
More than 90%, going from a day-zero kickoff to full production in nine days. The reduction came from summarizing repetitive log patterns rather than sampling or dropping data.

Did cutting log volume disrupt Envoy's engineers or break their dashboards?
No. Grepr automatically scans existing dashboards, alerts, and log-based metrics and adds them as exceptions, so the logs powering them keep flowing untouched. Engineers did not notice the change, and a follow-up survey returned a single minor comment.

Why didn't Datadog exclusion rules fix the cost problem?
They turned into a whack-a-mole exercise. The granularity was too fine to write a rule that killed the noise without risking a log the team might need during a P0, and costs crept back up within weeks as engineers shipped new features.

How does Grepr reduce log volume without losing data?
Grepr assigns a pattern to every message, passes rare, unique, and error logs through unchanged, and summarizes the noisy repetition. All raw data is still written to your own S3 bucket, so you can backfill any of it into Datadog in about thirty seconds when an investigation needs it.

Does summarizing logs affect compliance or data retention?
No. The raw data lives in your own S3 bucket, so you set your own retention policies and can delete whatever you need to. Envoy maintains SOC 2, HIPAA, CMMC, and ISO 27001, and reported no compliance issues.

How long did the rollout take?
Nine days from kickoff to full production. The staging setup took minutes, and Grepr's team ran one-on-one sessions to get Envoy's engineers up to speed.

Ready to reduce your observability TCO by 75%?
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..