Security teams have had the same conversation for decades. Too much data that’s too expensive to manage, so they don’t keep it all available. Then an incident hits, and the data they discarded turns out to be exactly what they needed.
“I’ve lived this exact scenario,” said Michael Cucchi, CMO & CPO of Hydrolix, a real-time data platform for internet-scale operations. “Earlier in my career, we were in the middle of an active investigation when we discovered we weren’t ingesting our audit logs. They’d been cut from the budget. Management didn’t think they were worth the line item because, in that moment, the storage cost looked like the risk. We ended up two weeks behind on the investigation, worried our core IP had walked out the door, rehydrating logs we should have had the whole time. That’s the moment the math flips on you. The cost of not having the data becomes bigger than the cost of storing it. But by then, you’re already behind.”
It’s a familiar scenario. One that the SIEM was supposed to fix. It provided a centralized correlation engine, full enterprise visibility, and one place to detect and investigate threats. While the concept was always right, the execution was expensive, hard to tune, and increasingly unable to keep pace with the environment it protects.
“The SIEM wasn’t wrong,” said Cucchi. “It was just built for a world that doesn’t exist anymore. It promised one place to detect and investigate everything, and for a while, that was true. But the model charges you at every step — ingest, process, store, retrieve — and that math only works if your data stays roughly the size it was when the pricing was designed.”
It didn’t. Data volumes kept growing, and instead of the SIEM scaling with them, teams started making silent tradeoffs: shorter retention windows, sampled data, logs quietly dropped because they were too expensive to keep.
The crux of the problem is that the data required for serious security analytics has grown exponentially, but SIEM architecture and cost models haven’t moved with it. For example, many platforms today can afford 30 to 90 days of hot data retention. For chatty traffic like XDR and endpoints, it can drop to seven days, which may not be enough to detect today’s slow, persistent threats. Sophisticated adversaries understand this. They operate on longer timelines, rotate identities specifically because they know what their targets are throwing away and how to skate below detection.
That was before AI came into the picture.
The cybersecurity industry has talked about AI for years, mostly in anomaly detection, alert triage, and accelerating investigation. All of which are useful, but still fundamentally human-paced. A system surfaces something, a human investigates, and a human decides. The feedback loop might be minutes. It might be hours. But a person was always in the middle.
That’s changing.
“We’re moving into an agent-to-agent economy, and in that world, decisions happen at machine speed, not human speed,” said Cucchi. “If an agent is deciding whether to isolate a host, revoke a credential, or let a transaction through, it has to make that call in the same window the threat is acting in, which is sub-second, not sub-minute.”
AI is moving into decision-making and has agency. A security analyst logging into their SIEM today increasingly sees a ranked, pre-triaged view. They don’t see 100 undifferentiated alerts. Instead, it’s a prioritized list where the system has already categorized severity and suggested a resolution. The analyst confirms, and the resolution happens in seconds. The old pattern, which included a manual investigation, some SQL query experience, maybe rehydrating data, and finally having the aha moment of mitigation. All of that now happens in the first few seconds of incident response, if it happens manually at all.
But the AI is only as accurate as the data you hand it, which is the next battleground in security: data completeness. An AI agent working from partial context misses telemetry, faces gaps in the forensic record, and has incomplete signal correlation. This causes the agent’s accuracy to drop. But give it a full context window and highly granular correlated signals with nothing missing, and its accuracy climbs.
“An AI agent investigating an incident doesn’t work off a hunch about which thirty days matter. It needs the full history to correlate a pattern that started months ago with something happening right now, and it needs to search all of it in seconds, not rehydrate a fraction of it in hours,” said Cucchi. “The organizations that kept everything hot and queryable are the ones whose agents can actually investigate. The ones who archived to save money are the ones whose agents hit a wall exactly when it counts most.”
The SIEM economics that worked at moderate data volumes are buckling now, especially when you bring in XDR data, endpoint telemetry, CDN signals, and WAF and API gateway logs. A large enterprise with hundreds of thousands of endpoints and distributed microservices can create a telemetry volume that is staggering.
Then you add agents. Agentic AI generates its own exhaust, which includes MCP visibility and governance, access logs, privilege events, and decision records.
“Ultimately, agents need to be treated as employees,” said Dr. Chase Cunningham (a.k.a. Dr. Zero Trust), a twenty-year veteran in comprehensive zero trust strategy consulting. “Every AI operating inside a company’s environment is also a zero-trust problem. Agents hallucinate. They make wrong choices. That’s not a knock on the technology; it’s a design constraint, and managing it requires data operators actually have.”
Another huge, current challenge in securing agents will require a new level of behavioral analytics. The UEBA (User and Entity Behavioral Analytics) that transformed the SIEM industry by enabling detection of abnormal employee and device behaviors will need to be adapted to do the same for AI agents. Is the action desired or malicious? Did it succeed or fail? What is the blast radius? What level of risk is there, and should I be signalling and mitigating? Agents are a whole new productivity but also an attack surface.
A new technology layer is now emerging. Analysts are starting to call it a security data fabric.
“The SIEM doesn’t have to be where all the data lives, but it has to be the place where all the data is reachable, in real time, at the moment of investigation, and at the moment an AI agent needs to act,” said Dr. Cunningham. “That takes a different underlying architecture. The architecture requires independent ingest scaling, independent compute, and object storage cheap enough to keep all data hot instead of archiving it somewhere where you’ll spend hours finding and rehydrating during an active incident.”
One last challenge with the emerging agent-to-agent economy is time. While AI-assisted humans still work in minutes and hours, AI agents work in seconds and sub-seconds. For businesses to compete, they’ll need to master and optimize for sub-second agent experiences. And in security, responders have sub-second windows to ensure the behavior is desired, abnormal, or malicious. So here ingest pipeline delays, data staging and indexing delays, and time to insight will need to innovate.
“This is the problem Hydrolix was built for,” said Cucchi. “Our company’s founders were generating massive telemetry volumes and paying unsustainable costs to manage them six years ago, so they designed a new data architecture for global-scale volume and real-time analytics.”
The same compression ratios and query speeds that made Hydrolix the platform of record for CDN observability apply directly to the agentic security data problem. The company has brought time-to-insight on more than a petabyte a day worth of data down to 5 seconds. Five seconds means catching the incident while it’s happening. It also means being able to inspect, assess, and secure agent-to-agent behaviors without interrupting the emerging business-critical workflows.
“We were ingesting nearly 200 terabytes of CDN log data over the course of the game, peaking at 17.4 gigabytes a second, and the queries still had to come back in under a second,” said Michael Cucchi, CMO at Hydrolix. “That wasn’t against a sample. It was against the full dataset because the moment you start sampling is the moment you miss the thing that actually breaks. Security teams are about to hit that same wall. An AI agent deciding whether to act in real time doesn’t get the luxury of querying 30 days and hoping the other eleven months don’t matter.”
A few things are worth accepting, even when they’re uncomfortable.
SIEM economics is creating a reckoning. The model where one vendor charges to ingest, process, store, forward, and retrieve data, and collects money at every step, doesn’t hold at the volumes security teams are managing now. The architecture is separating those functions, and the leaders who get ahead of that will most likely build something more durable.
Throwing data away can be a liability. Every log in cold storage is a detection and forensic gap waiting to surface at the worst moment. Every CDN signal discarded for being too noisy may be a web attack surface that is overlooked.
Agentic AI in the SOC is something that companies should prepare for now. The organizations that may actually benefit from it are the ones that have done the unglamorous work of building a stack that can support it technically. Making their data complete, correlated, and queryable in real time is table stakes.
Decades later, the conversation is the same, although with a faster clock: keep the data, get to it fast, and stay current with the architecture of the applications you’re protecting. The applications are moving faster than ever. The infrastructure can keep up, although the playbook that got the industry here may need to change.




