Back
Author
John Withers
Head of Marketing
LAST UPDATED
July 23, 2026
READING TIME
10 min read
Case Studies

How FOSSA Reduced Distributed Trace Volume by 95% Without Changing How Engineers Work

SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..

FOSSA provides a comprehensive solution for open-source license compliance, software bill of materials (SBOM), and vulnerability management, helping companies manage their open-source software usage effectively. Their SaaS platform runs across multiple services on Kubernetes, with Datadog as their observability platform of choice. It's a setup that gives the team strong visibility into how their systems are performing, and one that comes with a cost structure that scales directly with data volume.

The Problem: Unsustainable APM Costs

As FOSSA's microservices architecture scaled, distributed trace volume grew alongside it. Datadog's APM billing, which includes host-based components, became increasingly difficult to predict as FOSSA's clusters scaled during peak loads. But they needed to collect all their traces, because they didn't know which could be important, and configuring sampling for every service is error prone and tedious. The result: paying for data that provided little operational value.

The operational cost was just as real as the financial one. Engineers were spending time on observability cost reviews and quota discussions instead of building product. Controlling trace volume the traditional way would have meant instrumenting and re-instrumenting services by hand, a slow and labor-intensive process with no guarantee of maintaining the signal quality the team depended on.

According to Dave Bortz, VP Engineering at FOSSA, “After seeing what Grepr did to reduce our logs noise, extending it to traces was an easy call. We were seeing the same pattern: a lot of volume, most of it not particularly useful, and a Datadog billing model that scaled with every new host we spun up.”

Implementing Distributed Traces Reduction in Less than an Hour

Having already seen what Grepr's signal processing engine could do with log data, FOSSA expanded the implementation to cover distributed traces by simply updating their agents to send APM data to Grepr. Grepr ensured FOSSA still had 100% visibility into the RED metrics for their APM services, while heavily sampling “normal” spans. The FOSSA team had already configured APM span retention filters to keep any spans that they deemed critical; to ensure they didn’t lose operational visibility, they added a rule to Grepr to always send these spans to Datadog. This allowed FOSSA to migrate quickly to Prod with confidence, because they knew important spans were being retained. Ultimately, FOSSA had the same goal for traces as they had for logs: dramatically reduce what reaches Datadog while ensuring nothing operationally important gets lost.

Several capabilities made this work for FOSSA's trace environment specifically:

High-fidelity sampling that preserves what matters. Grepr's signal processing engine identifies and eliminates redundant, low-value traces while ensuring high-value signals pass through without modification. Traces containing errors or high-latency outliers are prioritized for ingestion into Datadog. Routine, successful traces that follow known patterns are the ones that get reduced. Two things worth calling out:

  1. Control: Users can decide what types of signals are “important” to them, allowing them to tune the signal processing engine.
  2. Simplicity: Users exercise this control at a single, platform layer, instead of needing to tune multiple configs.

Automatic detection of incomplete spans. FOSSA's instrumentation had cases where child spans were missing their root trace IDs, creating incomplete traces. Grepr identified these misconfigurations at the pipeline level, allowing for assembling traces before sampling decisions were made, offering cleaner, more complete trace data in Datadog.

SQL-based enrichment at the pipeline layer. Using Grepr's SQL operator, FOSSA was able to enrich spans in transit. To preserve Datadog's filtering and querying capabilities after host-level reduction, Grepr tagged spans with cluster name attributes, allowing the team to continue querying by cluster without any changes to how the application was instrumented. And because Datadog bills per unique host, in the past when FOSSA’s clusters would scale up in response to increased load, their Datadog bill would scale linearly. But now with Grepr, FOSSA’s ability to enrich spans ensures their host costs stay the same even as clusters scale up, instead shipping raw spans to their own S3 bucket at a fraction of the cost.

Service-specific configuration. The FOSSA team is able to configure which spans are always important, such as 500 error codes from the core service, and ensure they’re always sent to Datadog so their Dev team can investigate immediately.

The Results

FOSSA achieved a 90% reduction in distributed trace volume, consistent with what they had already seen on the log side. Datadog APM costs dropped sharply, and the scaling pressure from cluster growth became significantly easier to manage.

Critically, none of this required changes to how engineers instrument code or use Datadog day to day. Dashboards continued populating. Alerts continued firing. When engineers needed to investigate a specific trace, it was available in S3 and able to be backfilled into Datadog within seconds. The troubleshooting workflow the team had built around Datadog remained intact.

As Dave put it, “We wanted the same outcome we got with logs, and we got it. Fast implementation, no code changes, no dashboard rebuilds, no workflow disruptions. Our engineers kept working exactly the way they always had, and the costs came down. That's the result you're hoping for but don't always get.”

What's Next

FOSSA continues to expand their Grepr implementation across their environment. With both log and trace pipelines in place, the team now has a consistent approach to telemetry cost management that scales with the platform rather than against it.

Ready to reduce your observability TCO by 75%?
SHARE
Keep up with Grepr
Subscribe now for best practices, research reports, and more..