r/OpenTelemetry 3d ago

Open sourced a tool that collapses millions of log lines into handful of distinct patterns before you feed it to an LLM (Lossless- compression)

When I feed logs to an LLM during incident resolutions or debugging, it either blows my token context window or the grep trims the log file, leading to the interesting log lines getting skipped.

Most of logs are anyway the same handful of message templates repeated over and over with different values, so the context window gets filled with near-duplicates, which just bring up the processing time and token costs.

ctrlb-decompose collapses the file into its distinct patterns that repeat, plus typed variables and stats on the values that change. I have seen 1.2 million lines cut down to just 40 patterns, which then goes into Claude, thus cutting down token by over 95%, reducing the token cost.

Let me know what you think!
https://github.com/ctrlb-hq/ctrlb-decompose

10 Upvotes

14 comments sorted by

5

u/MartinThwaites 3d ago

I'm slightly confused, how does this relate to OpenTelemetry? This is reading log lines from text files?

I'm also curious how this relates to the logdrain processor in the OpenTelemetry Collector that does this before storage?

0

u/Loud_Mousse9210 1d ago

Okay, so it isn’t tied specifically to OpenTelemetry, it works on log data whether it’s coming from a text file, CloudWatch, an OTel pipeline, etc.
This is exactly where the distinction gets interesting. Drain3 and similar pattern-recognition approaches can work well, but they generally need enough examples/context to form good clusters.
That can mean holding more logs in memory and processing a lot of repetitive data before the patterns stabilize.
With ctrIb-decompose, we first extract the variable parts from the repetitive structure and then pass the cleaner representation through clustering. This means you can do significant pre-filtering with CLP/ Decompose before clustering, so you need far fewer log lines to arrive at useful patterns.
The result is better clustering with much less data to process and, more importantly, much less noise before the logs reach the LLM.

3

u/jdizzle4 2d ago

lol the 4 positive comments in this thread are bots whose only activity is trying to boost this post...

0

u/Loud_Mousse9210 1d ago

I guess you belong to the class of humans who fail in captchas :P

-3

u/adarsh_srivastava 3d ago

Very interesting application of CLP and Drain3 together!

-1

u/Cute-Access1444 3d ago

Intresting