Back
Author
John Withers
Head of Marketing
LAST UPDATED
September 28, 2026
Description
Datadog exclusion filters look like a cost saver, but they trade a lower monthly index bill for two costs you pay later: the never-ending work of maintaining the rules, and the slow, expensive rehydration you hit exactly when an incident needs the logs back.
Engineering Guides

The Hidden Cost of Datadog Exclusion Filters (and a Better Way to Cut Log Costs)

Blog post featured image
IN THIS ARTICLE
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..

If you run Datadog at any real scale, someone on your team is maintaining exclusion filters. It is the standard way to keep the indexing bill under control: decide which logs are noise, write a rule to keep them out of the index, and watch the number come down. Then a new noisy service ships, you write another rule, and so on and so forth. Never-ending, steady, unglamorous work. 

When most folks have conversations about observability cost, they stop at the obvious number: the monetary cost of the telemetry data they index and ingest. And yes, while that’s real money, it’s only a slice of the pie. Rarely is the rest of it counted, which is the time your team spends deciding what to exclude, and the time (and cost) involved in getting logs back when you actually need them.

Exclusion filters are a dangerous game

Exclusion filters are a bet that a given class of logs will not matter later. You’re probably right most of the time, because you have a sense of historical needs. The problem is that there will be times you’re wrong. It’s inevitable. And these are the moments where you absolutely don’t want to be up the creek (an incident) without a paddle (your data).

The common scenario

An incident happens, and the dev team is asking for logs from a service you excluded three months ago. You’ve never needed them in the past, so you decided to filter them out to save money. The logs can’t be found in the index, so you rehydrate them from your archive. In Datadog, that means waiting and paying real money. The rehydration will run against your cold storage (ie, slow) and is priced per million events rehydrated. The scan is sized by the whole time window you select, not the handful of logs you’re actually looking for, because the query filters are applied after the data is pulled. All that to say, during incidents, the excluded logs are the slowest and most expensive ones to reach. And that is the hidden cost of exclusions. You might save on the index each month, but you pay it back (plus some) when you end up needing it and under the most pressure.

What’s a better approach?

Grepr’s ML-powered, signal processing engine automatically identifies log patterns, trace signatures, and metric trends dynamically, instead of relying on static rules patterns. So you can send all of your telemetry data to Grepr, and it will automatically reduce the volume by summarizing repetitive patterns rather than excluding them. Take the case where your system emits 10,000 identical logs: Grepr would identify these logs as belonging to a single pattern and would send a single log to Datadog with some additional statistics, including a count of the number of logs, and this would repeat every two minutes. Unique messages (ie, novel patterns) pass straight through to Datadog. There are no lists of exclusion rules to manage, and you’re not playing a game of roulette trying to decide which ones you are willing to lose.

So back to the scenario. Say you need to analyze logs that Grepr had identified as “noise.” Every raw log Grepr receives is written to your own low-cost object storage, so it stays queryable using the same query language your team already uses. There’s no rehydration process or need to spend money to access the data. If something you reduced ends up being the data you need, a dev queries it in the data lake and pulls it back into the observability tool in real time.

Now, you no longer need to make tradeoffs or roll the dice with exclusion filters. You keep the cost savings you were after when you made the filters, you no longer need to build and maintain manual rules, and you stop paying additional money to Datadog while waiting for a slow rehydration process every time you need the data back. It’s your data, and you should be able to query it whenever you want.

But Datadog has Flex Logs, can’t I just use that?

‍Maybe. Datadog Flex Logs decouples storage from query compute, so you can keep high-volume logs queryable for up to 15 months at a lower storage cost, then pay a committed compute spend. This is a reasonable option for teams that want cheaper long-term retention of logs they occasionally need.

But here’s the difference, and why I say maybe.

  1. Flex keeps your data inside Datadog. Grepr lets you keep it in your own S3 bucket. You own it under your own retention and access policies. 
  1. Flex is a storage and query tier, not a reduction. You still pay to ingest every GB, then pay a fixed compute cost every month with a minimum size. Grepr summarizes the volume before it ever hits storage, so ingest and index come down too.
  1. Flex doesn’t support Log Monitors or Watchdog, so logs in Flex cannot alert the same way.

  2. Exclusion filters! Flex only changes where some logs live and how you pay for them, but you’re still on that rule-maintenance treadmill. With Grepr the exclusion rule writing stops completely.

But hey, maybe you love Flex, so let me add this: Flex and Grepr can co-exist. The point is that neither the index bill nor the rehydration penalty has to be the price of keeping costs down. 

Here’s the takeaway

Exclusion filters trade a lower monthly index bill for two costs you ultimately pay later: the ongoing work of maintaining the rules, and the time and money that goes into rehydrating logs when you inevitably need them. When you point your logs at Grepr, you stop writing the rules, the volume still comes down through summarization, and every raw log stays queryable in your own storage in your own query language. I might be biased, but that does sound like a much better solution, does it not?

Learn how Envoy cut Datadog costs by using Grepr here.

‍

FAQ

Do Datadog exclusion filters actually save money?

‍They cut your monthly index bill, but not the hidden costs: the ongoing work of maintaining the rules, and the time and money to get excluded logs back when you need them.

How much does Datadog rehydration cost, and how long does it take?

‍It's priced per million events and runs against cold storage, so it's slow. The scan is sized by the whole time window you pick, not just the logs you want, so you pay to scan the full range. Check Datadog's pricing page for current rates.

How is Grepr different from exclusion filters?

‍Grepr reduces volume by summarizing repetitive patterns instead of dropping them. Ten thousand identical logs become one summarized log with a count, while unique messages pass straight through, and the raw data stays queryable in your own storage. No rules to build or maintain, and no rehydration.

Is Datadog Flex Logs the same thing?

‍No. Flex is cheaper long-term storage, but it keeps data inside Datadog, still charges you to ingest every GB, doesn't support Log Monitors or Watchdog, and doesn't end the exclusion-filter treadmill. Grepr keeps data in your own S3 and reduces volume before it lands.

‍

‍

Ready to reduce your observability TCO by 75%?
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..