A foundation for observability
As AI agents take on more of our software, the evidence of what they do matters more, not less. We started Logpacer to keep that evidence: store more of it affordably, make it searchable fast, and stop asking operators to write parsers for the privilege of looking at their own logs. This is where the project began, what it took to get there, and why we're still building it.
Published
And why we’re building Logpacer
We believe it should be easy and straightforward to collect your logs, keep them, and search them when you need to understand what happened. All of them - not just the ones you decided, beforehand, that you could afford to keep.
Today, that is harder than it should be. Storage costs push you to trim earlier and earlier, and the machinery between “this software emits logs” and “I can search them” is complex enough to feel a little absurd. You can make an informed choice about which logs to keep, certainly. You can’t make it with knowledge of the incident you haven’t had yet.
When we started Logpacer, we kept coming back to that. We expected the amount of log data to grow, and we wanted to keep more of it without making storage prohibitively expensive or searching it impractical. Those requirements had to hold together. A wonderfully compressed archive that you can’t use when you need it isn’t much of an answer.
The industry data since then gives us little reason to expect that pressure to disappear. In its report covering April 2024 to April 2025, Cribl reported growth of more than 60% in each of three core telemetry inputs: syslog, TCP, and TCP JSON. That is a view of activity within its own Cribl.Cloud customer base, not a universal growth rate for the world’s logs, but it is a concrete indication of the pressure in systems already being operated.[1]
And now we’re giving software more work to do on our behalf. In June 2025, Gartner forecast that 33% of enterprise software applications would include agentic AI by 2028, up from less than 1% in 2024.[2] That is an adoption forecast, not a prediction of log volume. Our concern is what those applications will require us to understand: a single request can become a sequence of model calls, tool invocations, retries, and actions across other services. When something goes wrong, the final answer may be the least interesting part of what happened.
We believe your capacity to examine and investigate what happens on your infrastructure is now more important than ever.
So we started building a database
Building your own database gives you considerable freedom - including the freedom to discover exactly how much work you’ve volunteered for.
We’ve doubted that decision plenty of times. But we kept returning to the same constraint: efficient storage and useful search couldn’t be things we hoped to reconcile afterward. They had to shape what we built.
Logs gave us something to work with. Even when they aren’t emitted in a neatly structured format, there is often an implicit structure in how they’re written. If we could extract it and organize the data into columns, we could take advantage of that in both compression and search. We developed our storage format and query approach around those observations, gradually refining what became Sublog.
Of course, “extract the structure” is doing a suspicious amount of work in that explanation.
Logs don’t politely stick to one format, one event per line, ready for someone to put them into columns. There are exceptions spread across multiple lines, different formats wedged into the same output, and logs that wrap other logs. You work out how to read the outside, and discover that the inside has opinions of its own.
We were writing regular expressions, trying template-extraction approaches, and spending too much time resolving what the data actually meant. This was how we ended up building the analyzer: the database needed structure, and making that structure somebody else’s responsibility wasn’t going to remove the work. It would just determine who had to do it.
The logs still had to get there
The analyzer wasn’t the end of it. We also had to collect the data and transport it, which brought us into configuring OpenTelemetry collectors and setting up the machinery needed to manage delivery.
Each task was reasonable enough in isolation. Of course a collector needs configuration. Of course transport needs managing. But put them together - and remember that you originally wanted to look at some logs - and the arrangement starts to feel a little unreasonable.
We were slowly getting bitten by all of it. Not one spectacular problem that explained the whole difficulty, but the thousand paper cuts of getting from “this software emits logs” to “I can actually use them.” Each decision led to another thing we needed to understand, configure, or take responsibility for. What had begun with a storage format was drawing us into much more of the journey than we had set out to build.
There are capable tools for these individual jobs. That wasn’t the objection. What bothered us was how much work remained between them, and how reliably that work ended up with the person who was supposed to benefit from the whole arrangement.
A problem hasn’t disappeared just because it has moved into a configuration file they now own.
Couldn’t we just click the logs we want?
This became a difficult question to leave alone. Wouldn’t it be wonderful if you could just choose the logs you wanted to collect, and let the system turn that choice into a working setup?
Not learn the collector’s configuration language first. Not write the parser that makes the data useful afterward. Choose the sources you care about, without inheriting every implementation detail required to collect them.
Against everything we were dealing with, that sounded almost utopian. Which is a strange thing to say about selecting some logs.
There would still be plenty of complexity underneath. We knew that because we kept finding more of it. But did all of it really need to become something the person using the system had to learn and maintain? Choosing what to collect is their decision. Requiring them to become proficient in the collection machinery before they can make that choice is another matter.
We’ve come to think of it as compressing the complexity. The work still has to happen. The difficult cases don’t disappear because there is a simpler interface. The ambition is to take responsibility for more of that work, so using the system doesn’t require understanding all of it first.
That ambition grew out of the engineering. We didn’t start with a plan to build every part of this journey. We started with a constraint about storage and search, and kept encountering reasons why solving that part alone wasn’t enough.
The same work is also taking us toward greater visibility and control over traffic to AI services, closer to where it originates. More on that when it’s ready.
This series is about the decisions that brought us here: why we built the database, what it took to make real logs useful to it, how collection and transport became part of the problem, and what happens when formats change or delivery fails. We’ll explain the tradeoffs, including where the apparently simpler approach turned out to leave quite a lot of work unfinished.
What that looks like today
Today, Logpacer handles application logs, host logs, metrics, and traces. Selecting sources and getting searchable data isn’t only an ambition: in the ten-minute demo on our website, we go from a clean host to collected, structured, searchable logs, without writing a parser. That’s one demonstrated setup; environments with stronger opinions of their own may take longer.
Logpacer also provides native CLI and MCP interfaces, so agents can discover sources, inspect schemas, and query the collected data directly. The evidence is available to people and agents alike, without requiring a human to copy results between tools.
We keep the logs we collect verbatim, so the original record remains available when the next question isn’t the one you expected. Compression makes keeping that evidence practical; making it searchable makes it useful.
That’s what compressing the complexity means to us. There is plenty of machinery underneath. You shouldn’t have to assemble it before you can find out what happened.
Sources
[1] Cribl - Telemetry trends: Insights unveiled
[2] Gartner - June 2025 agentic AI forecasts