Can AI Fix Weak Security Data?

Can AI Fix Weak Security Data?

AI can help analysts move faster, but reliable outcomes still depend on complete telemetry, connected context, and detection logic strong enough to support the decisions that follow. 

Zubair Chowgale, Director, Sales Engineering, Securonix 

 

There is a kind of AI demo every security team has seen by now. An analyst asks a question, the system understands what they mean, searches across logs, connects the alert to a user, an endpoint, a cloud account, and a few recent events, then produces a clean summary of what likely happened. It looks effortless because the environment behind the demo is usually clean enough for the story to make sense.

Most production SOCs are messier. Endpoint data may tell one part of the story while identity data sits somewhere else. Cloud logs may arrive in a different schema, with different retention rules and a different idea of what a user or workload even means. Asset context may be stale. Application owners may have changed. Some logs may be missing the fields needed for correlation, and the one person who knows which source is reliable may be on vacation when the investigation starts. AI can move quickly through that environment, but speed loses value when the evidence underneath it is incomplete or disconnected.

The model may look ready. The workflow may look ready. The data foundation often is not.

What Happens When AI Meets Bad Telemetry?

Security teams have always known that telemetry quality matters, but AI makes weak telemetry much harder to ignore. A human analyst can often compensate for messy data because experience fills in the gaps. They know which integration drops fields, which cloud account has inconsistent tags, which identity source is reliable, and which endpoint agent tends to disappear at the worst possible moment. That knowledge rarely lives in the platform. It lives in people.

AI brings a different kind of limitation. It can summarize what it can see, while gaps in the evidence may stay hidden unless the system has been designed to expose them. A partial timeline may look complete when the missing data never arrives. Events that share a field may look connected even when the relationship is meaningless. A confident explanation may rest on inputs that an experienced analyst would question immediately. The concern is not that AI gets everything wrong. The concern is that a polished answer can arrive before the organization has proven the data is strong enough to support it.

A 2025 empirical study of LLM use in SOCs gives a useful window into how analysts are actually using these systems. Researchers studied 3,090 analyst queries from 45 SOC analysts over 10 months and found that LLMs were often used as on-demand aids for sensemaking and context-building, not as replacements for analyst judgment. Many of those interactions involved interpreting low-level telemetry, including commands, and helping analysts work through technical context. That finding matters because it places AI exactly where telemetry quality becomes visible: inside the practical work of understanding what an event really means.

For many organizations, the first lesson from AI in the SOC may be uncomfortable. The issue is usually a lack of connected data that analysts and AI systems can trust at the moment a decision needs to be made.

Why Does Security Context Matter More Than Volume?

Most SOCs have plenty of events. Meaning is harder to come by.

A failed login becomes useful only when it is connected to who attempted it, where it came from, whether the device was known, whether the user normally authenticates that way, what privileges were involved, and what happened before and after. A suspicious process needs the parent process, the user context, the asset value, the vulnerability state, and the surrounding behavior before it becomes something an analyst can trust. A cloud permission change may be harmless or urgent depending on who made it, what resource was affected, and whether that privilege is normal for the workload.

AI can help analysts move through evidence faster, but relationships still need to exist in the security data before a system can reason over them reliably. If identities are not resolved consistently, if cloud resources are poorly tagged, if assets are not mapped to business owners, or if detection logic relies on fields that are often missing, the same weaknesses will surface in different forms. Sometimes the answer will be incomplete. Sometimes the summary will sound polished but leave an analyst with the uncomfortable feeling that something important is missing.

This is why AI readiness in the SOC starts long before anyone deploys an assistant into an investigation workflow. It starts with the ordinary, tedious questions that determine whether data can support a decision. Which log sources are authoritative? Which identities can be resolved across endpoint, cloud, SaaS, and network activity? Which alerts depend on enrichment that arrives too late or too inconsistently? Which investigations still rely on tribal knowledge because the workflow never preserved the context analysts needed?

Is Security Data Architecture Part of AI Strategy?

The organizations that see the earliest value from AI in security operations will probably be the ones that already understand their environment well enough for AI to reason over it.

That starts with data architecture. Telemetry needs consistent schemas. Identities need to resolve across systems. Assets need ownership, criticality, and business context. Cloud, endpoint, network, and SaaS events need enough shared structure that an investigation can move from one source to another without losing meaning. Detection logic needs to be validated, versioned, and refined as the environment changes.

Recent research on canonical security telemetry makes this point in more technical terms. The authors argue that AI-driven cyber detection can struggle when telemetry remains fragmented, event-centric, and inconsistent across environments, and they propose entity-based and relationship-aware telemetry as a stronger foundation for AI-native detection. That maps directly to what practitioners see in production. Security data becomes more useful when it can describe persistent identities, assets, relationships, and behavior over time rather than isolated events scattered across tools.

Security leaders should treat AI adoption as a security data architecture project, not only a tool rollout. A model may be able to summarize an alert, but the useful question is whether the organization has built the telemetry, enrichment, and governance needed for that summary to be trusted.

Can AI Improve Detection Without Better Detection Engineering?

Detection engineering becomes more important as AI enters the SOC.

Natural language interfaces will make it easier for analysts to ask questions. AI-assisted triage will make it easier to summarize activity. Agentic workflows will make it easier to route evidence, generate hypotheses, and recommend next steps under human supervision. Strong detection logic still determines the quality of the signal those workflows depend on.

A poorly tuned rule remains noisy even when the summary is clear. Missing fields still create blind spots. Inconsistent naming still breaks correlation. Detection logic that cannot be tested, explained, or improved will limit the value of any AI-assisted workflow built on top of it.

This is why mature SOCs are becoming more engineering-driven. Detection-as-code, version control, schema management, validation pipelines, and repeatable enrichment are becoming practical necessities rather than maturity-model luxuries. AI can support that work, but it depends on a feedback loop where investigations improve detections and detections improve future investigations.

Human oversight matters here as well. The SOC study found that analysts used LLMs as aids for sense-making and context-building while preserving analyst decision authority. That distinction is important. The best use of AI in security operations is not blind delegation. It is faster context, clearer evidence, better hypotheses, and human review that can trace the reasoning back to the data.

A faster SOC only helps when the people inside it can still understand why a decision is being made.

What Should Security Teams Fix Before Scaling AI?

A practical AI readiness test begins with the investigations analysts already run every day.

Where do they lose time? Which context has to be found manually? Which fields are missing when they matter most? Which systems describe the same identity differently? Which detections create noise because they lack business context? Which escalation depends on one analyst who happens to know how the environment really works?

Those answers show where AI will struggle before it is deployed broadly. They also show where AI can become useful once the foundation improves. Better identity resolution makes summaries more reliable. Better asset context makes prioritization more meaningful. Better cloud telemetry makes investigations more complete. Better detection engineering gives AI stronger signals to reason over. Better evidence trails make human review faster and safer.

AI will change security operations by helping analysts summarize evidence, accelerate triage, generate hypotheses, and remove repetitive investigation work. It will also expose the places where the data foundation is fragmented, inconsistent, or incomplete.

The practical point is straightforward. Weak security data becomes much harder to ignore once AI enters the workflow, because every assisted investigation depends on the same telemetry, context, and detection logic the SOC already trusts today.