CISO Playbook: Putting Claude to work in security operations

August 19, 2026

Written by

Itai Tevet

This playbook is for security leaders who know AI belongs in the SOC, but need a practical model for where it actually fits. It’s written for CISOs, SOC leaders, detection engineers, and security teams dealing with alert volume, manual triage, reporting drag, and pressure to justify AI investment. The goal is to separate what should be handled autonomously from what should stay human-led with a copilot like Claude. If you’re trying to move from scattered AI experiments to a repeatable SecOps operating model, this guide is for you.

1. Executive summary

Most security leaders already accept that AI agents like Claude can change how their team works. The harder question, and the one this guide answers, is where to start.

The FOMO is justified. Coding and agent platforms have moved past the demo stage into real SecOps work including alert triage, detection engineering, reporting and more. But a lot of teams bolt a chatbot onto a console, or wire an agent straight into their detection tools, get a mediocre result, and decide AI isn't ready yet.

A lot of teams bolt a chatbot onto a console, or wire an agent straight into their detection tools, get a mediocre result, and decide AI isn't ready yet.

Here is the model this guide argues for. Think of two halves of one brain, not two rival products. One half is autonomous. It’s an AI SOC platform that takes the repetitive, around-the-clock work like alert triage off the team completely. The other half is copilot tooling like Claude and its peers, which make the judgment work (that actually requires human involvement) faster without taking it away from analysts. Like the two hemispheres of a brain, they do different jobs, and the system only works well when both are present.

What you'll walk away with:

  • A model for where AI actually fits in your SOC for both the autonomous half and the copilot half, and why confusing the two is the most common and most expensive mistake.
  • Which use cases the copilot pays off on first, each with a before/after and the catch, so you can judge the value yourself.
  • A phased rollout that reaches real value in as quickly as three months.
  • The metrics to defend the spend to your board.
  • The mistakes that quietly kill adoption and how to sidestep them.

NOTE: This guide is not about adopting AI securely. Data handling, model risk, and governance all matter, but they belong in a different document. For that, start with the NIST AI Risk Management Framework for governance and risk, and the OWASP GenAI Security Project (Top 10 for LLM Applications, plus the 2026 Top 10 for Agentic Applications) for the engineering-level risks.

2. Why now: defense against machine-speed attackers

The economics of attack just changed, and defense has to change with them. Attackers aren't limited by human hands at a keyboard anymore. Agent-driven tooling lets them probe and phish at machine speed and machine scale, and that's a volume a human-staffed SOC was never built to absorb.

The standard SOC answer to volume is prioritization. Work the loudest alerts, triage whatever the shift can reach, and quietly let a long tail of lower-severity signals go unexamined. When attack volume was bounded by human effort, that was uncomfortable but survivable. Against machine-speed adversaries, that unexamined tail is exactly where intrusions live.

The idea running through this whole playbook is that AI executes and humans supervise. Not AI replacing analysts. AI becoming the 24/7 capacity that looks at everything, so analyst judgment gets spent where it counts. A capacity-bound SOC normally can't have complete coverage, consistent depth, and a working feedback loop at the same time.

This playbook shows how to get:

  • Every alert investigated
  • The same quality at 3 a.m. as at 3 p.m.
  • Every human decision feeding back into detections and investigations

This isn't a scare story. It's an opportunity that happens to arrive right when the threat model demands it.

3. What goes wrong: the two ways teams misapply AI

The gap in most security orgs isn't missing AI. It's misapplied AI. Two failure modes show up over and over. The first is the bolt-on chatbot: handy for the odd question, but cut off from the team's tools and context, so it never touches the real work. The second is the opposite mistake, wiring an autonomous agent straight into detection sources and expecting it to run the SOC. General-purpose agents wait for a prompt. They don't sustain high-volume, around-the-clock investigation on their own.

Both fail for the same reason. They ignore which layer the work belongs to. So before anything else, figure out which layer you're operating at today. Most teams have never actually named it.

Where are you on the curve?

The ladder works two ways. It tells you where you stand, and it shows you the next rung. The rest of this playbook is how you climb.

4. What good looks like: the AI SecOps stack

It helps to picture modern security operations as a computing stack with three layers. Once you see the layers, the "complementary, not competing" point stops being a slogan and starts being obvious.

Detection (the hardware) EDR, SIEM, cloud, and identity tools each watch their own slice and alert on it. On their own, they're the reason a busy SOC drowns in thousands of disconnected signals a day.

The AI SOC (the operating system) This is the layer that makes the hardware usable. It investigates alerts on its own, ties findings together across tools, and holds the org's memory: the detection rules, the case history, the forensic context, the integrations.

AI platforms (the applications) Claude, Codex, and the rest are where humans do supervised work such as tuning, hunting, reporting, and the calls that need a person.

What connects the middle and top layers is the Model Context Protocol (MCP) of the AI SOC. It's what lets Claude work off the AI SOC's normalized cases, collected evidence, and live SIEM/EDR queries with one correlated picture instead of a dozen raw feeds it has to assemble itself.

For more details, see "The other half of the AI SOC: Intezer, now inside your AI workspace".

The stack lines up with a simple split that drives everything downstream. Some work you want gone entirely: repetitive, high-volume, around the clock, with alert triage as the obvious example. That belongs to the autonomous layer. Other work you want to keep but do faster: investigation, hunting, decisions that need judgment. That belongs to the copilot, with a human in the loop.

One requirement is worth stating plainly. If you bring in an autonomous layer, it has to investigate every alert, including the low- and medium-severity ones, not just the loud few. A layer that only reads high-severity alerts rebuilds the exact blind spot you were trying to close, and it wrecks the economics. Full coverage is what earns it your trust.

A few principles hold the model together. The layers compound: every tuning rule an analyst writes in the copilot makes the autonomous layer a little smarter, so there are fewer escalations next time. Claude and an AI SOC shouldn’t be seen as rival products; they sit at different heights in one stack, and confusing them is the most common strategic error. Context is what makes or breaks the whole thing, because an AI platform is only as good as the detection logic, case history, and live integrations you put underneath it. And the posture throughout is supervision rather than replacement. AI executes, humans supervise, and you build for a real human in the loop instead of automation theater.

5. The use cases: Claude at work in the SOC

These are the jobs where a copilot earns its place fastest, work you could put Claude on tomorrow morning and feel the difference by the end of the week. They deliberately sit above the autonomous layer: the AI SOC already triages the flood of alerts on its own (see previous section), so what's left here is the work that still needs a human's hands and judgment, just done far faster. Each one follows the same shape, the job, what it costs today, what changes with Claude, the payoff, and the catch, so you can weigh the value for yourself.

5.1 Working an escalation: from a cold hand-off to a confident call

The job to be done: The autonomous layer has already done the heavy lifting. It investigated the alert at forensic depth and escalated the ~2% where the technical facts are settled but the decision turns on business context no security tool can see. Take an impossible-travel alert: the AI SOC confirmed the device is managed, MFA passed, and the login fits the user's baseline. What it can't know is whether this person was actually authorized to travel. That call is yours and to make it, you still have to gather context the security stack never holds.

Before: Even with a clean hand-off, the analyst pivots across tools the AI SOC doesn't touch to make the business call: pinging the user's manager on Slack, checking the shared travel calendar, scanning email for an "I'll be in Berlin next week" thread, cross-referencing the ticketing system. The context is scattered, so a decision that should take two minutes takes twenty, and how thorough it is depends on who's on shift.

With Claude: The analyst picks up the escalation in the copilot, which inherits the full investigation trail from the AI SOC and reaches the context the security tools never see, e.g. email, Slack, calendar, ticket history, all in a single pass.

> Case #4821 was escalated as a possible impossible-travel compromise. The AI SOC has
> already confirmed the device is managed and MFA passed. Check our Slack, the shared
> travel calendar, and this user's recent email for any sign the trip was authorized,
> and summarize what you find with a link to each source. Flag anything you can't confirm.

What comes back isn't a fresh investigation. That part is done. It's the business context assembled in one place: the Slack thread where the trip was mentioned, the matching travel-approval ticket, the calendar entry, and anything it couldn't confirm flagged out in the open.

The payoff

The cases that genuinely need a human get one faster, with the gathering already done. The analyst spends their time on the judgment call, the part that actually needs them, instead of reassembling context from five tools. And the trail reads the same at 3 a.m. as it does at 3 p.m.

The catch

The copilot is only as good as the context you connect it to, and the business call is still the analyst's to make and own. Treat what comes back as input to a decision, never the decision itself. This is the human-in-the-loop slice, one case at a time. Investigating every alert around the clock is the autonomous layer's job (Sections 4 and 5.5), not an analyst feeding alerts into a chat window all night.

Before (manual)With Claude (copilot)Time to decision20+ min of pivotingA few minutes of reviewContext reachLimited to whatever you manually checkEmail, Slack, calendar, tickets in one passConsistencyVaries by analyst and shiftSame depth every time, especially if coupled with prebuilt skills.Analyst roleRe-gathering scattered contextReviewing and deciding

5.2 Detection-as-code: write, test, and tune in one loop

The job to be done: A new technique shows up, in a threat report, an incident, or a gap you noticed, and you need a detection for it: written in your rule language, tested against real data, checked for false positives and evasion, and shipped.

Before: Detection engineering is a context-heavy slog. An engineer hand-writes the rule in whatever the stack speaks, Sigma, KQL, SPL, runs it against historical data, eyeballs the hits for noise, tweaks, and repeats. Translating a rule from one SIEM's language to another during a migration is a multi-day job on its own. Good detection engineers are scarce, so the backlog just grows.

With Claude: The engineer describes the behavior, lets Claude draft the rule, and then drives a tight loop in the copilot. This is far smoother when Claude sits on top of an AI SOC layer: through MCP it already has the schema context and a safe, consistent way to write and test queries against your data, unlike a raw, hand-wired connection to each SIEM or EDR, where the query dialects and access are inconsistent and easy to get wrong.

> Write a Sigma rule for OneDrive being used to exfiltrate data via shared links.
> Then convert it to our Sentinel KQL, test the logic against the sample events attached,
> list the false positives you'd expect, and propose a narrowing condition for each.

Claude produces the rule, the cross-SIEM translation, a set of test cases, and a candidate tuning, all reviewable in one place. The engineer stays in the driver's seat. The typing, the translating, and the first-pass testing collapse into minutes.

The payoff

The backlog burns down faster, rule coverage widens, and a SIEM migration stops being a quarter-long rewrite. Junior engineers punch above their level because the syntax isn't the bottleneck anymore. Judgment is. (The backlog is structural: in the detection engineering research from ESG and Anvilogic, 86% of security professionals said it takes a week or more to ship a single new detection and CardinalOps' annual SIEM research pegs typical SIEM coverage at roughly one in five known ATT&CK techniques. The write–translate–test grind is exactly what the copilot collapses.)

The catch

A generated rule that looks right can be quietly wrong: too broad, too narrow, or blind to an obvious evasion. Every rule still gets reviewed and tested against real data before it ships. Claude speeds up the loop; it doesn't replace the engineer's sign-off. Feed it bad sample data and you'll get confident, useless rules.

5.3 Malware and phishing analysis: a verdict with its reasoning, in minutes

The job to be done: A suspicious script, attachment, or URL lands, from a phishing report or an alert, and someone needs a verdict: what is it, what does it do, and is it a real threat?

Before: This is specialist work. An analyst deobfuscates the script by hand, traces what it does, detonates the sample in a sandbox, and pieces together a verdict. It's slow, it leans on reverse-engineering skill that's rare on most teams, and phishing reports stack up faster than anyone can work them.

With Claude: Claude reads obfuscated scripts and walks through them step by step, pulls out indicators, and explains the behavior in plain language, which turns a reverse-engineering task into a review. For automated verdicts at real volume, this is where the autonomous layer earns its keep. A forensic AI SOC platform such as Intezer analyzes the artifact, returns a verdict backed by evidence, and hands the analyst a complete picture instead of raw sandbox output. The copilot is where a human then digs into the edge cases.

The payoff

Phishing and malware triage that used to need a specialist now gets a fast, evidence-backed first pass, so the queue moves and the specialists spend their time on what's genuinely novel.

The catch

Dynamic behavior still needs a sandbox, and a confident summary of a novel sample can miss something real. This is exactly where a good AI SOC layer earns its keep: through MCP it hands Claude real analysis tools such as sandbox detonation, forensic artifact analysis, verdicts backed by evidence, etc. so the work stays safe, consistent, and fast, and doesn't burn a pile of tokens re-deriving what a purpose-built engine already knows. Use Claude to understand and triage quickly, then confirm true verdicts with actual analysis. Never run unknown samples outside a contained environment.

5.4 Reporting: from a day of assembly to a one-prompt draft

The job to be done: Reporting never stops, and it eats senior time: incident reports, executive and board summaries, posture and visibility roll-ups, and the weekly metrics nobody enjoys assembling.

Before: Someone pulls the case timeline by hand, expands it across affected users, devices, and IPs, chases stakeholders for context, and formats the whole thing into something an executive will actually read. A single incident report can eat most of a day, and the quality swings with who wrote it and how tired they were.

With Claude: Reporting becomes a draft-and-review task. For an incident: pull the case timeline, expand across the affected entities, write an executive summary with the technical detail underneath, and export a styled PDF, all from one prompt. The same pattern covers the recurring stuff too, like posture summaries, board narratives, and metrics digests built straight from the case and detection data.

> Draft an incident report for case #4821: timeline, affected users/devices/IPs, root
> cause, actions taken, and a 5-sentence executive summary up top. Keep it factual and
> cite the evidence for each claim.

The payoff

Senior analysts get hours back per report, executives get consistent formatting and a narrative they can trust, and reporting stops being the thing that slips when the team is slammed. (The burden is real and still manual: per the SANS 2025 SOC Survey, 79% of SOCs run 24/7 yet 69% still rely on manual reporting — senior hours that draft-and-review hands back.)

The catch

A report is only as truthful as its source data, and a fluent summary can launder a bad assumption into something that sounds authoritative. The person who signs the report still owns its claims. Claude drafts; a human verifies before it goes to a board or a regulator.

5.5 The handoff: where the two layers meet

The job to be done: Out of thousands of daily alerts, a handful genuinely need a human. The job is to get that handful in front of an analyst with everything done except the judgment.

Before: Either everything routes to humans and they drown, or crude severity filters auto-close the quiet alerts and real threats slip through. Neither one is coverage.

With the stack working together: The autonomous layer investigates every alert, around the clock, and closes the vast majority on its own. It escalates only the cases that genuinely need human interaction or a decision, where the technical facts are settled but the call needs business judgment. The analyst picks those up in the copilot and inherits the full investigation trail, plus the context security tools never see such as email, Slack, ticket history, etc. They decide, and their tuning flows back to make tomorrow's autonomous triage sharper.

The autonomous and human layers working as one continuous loop with a handoff between them

The payoff

Complete coverage and a sane analyst workload, out of one system. Humans only see what deserves a human, and every decision they make compounds (Section 4).

The catch

This only holds up if the autonomous layer really does cover all severities and hands off with its evidence. A black-box escalation that just says "look at this," without the trail, only moves the manual work around. Demand transparency at the boundary.

Where else teams put the copilot: The same pattern of human judgment and AI legwork, covers more than these four. Teams also use it for hunting hypotheses run across normalized telemetry, purple-team and tabletop exercise design, vulnerability triage against real asset exposure, compliance evidence gathering, and post-incident reviews drafted straight from the case record. If it's judgment work wrapped in assembly work, it's a candidate.

6. Choosing your copilot platform

The copilot layer isn't one product, it's a category, and the platforms mostly differ in how you work with them. You'll probably standardize on one as your through-line and let specialists use what fits them. At Intezer, by the way, we use Claude, though in a space moving this fast, that can change. One caveat: these tools are converging quickly. Their capabilities overlap heavily, and the honest differences are more about working style and where each is strongest today than hard feature walls. Many teams end up using more than one.

Copilot platforms at a glance

ClaudeCodexCursorHow you work with itChat, CLI, and agentic terminal (Claude Code); browser controlCloud/CLI agent you hand async tasks toAI-native IDE (visual editor)Where it shinesDeep, multi-step agentic work across tools; strong real-world problem-solvingAutonomous, parallel coding tasks; token-efficientFast, interactive in-editor codingExtensibilityBroad MCP ecosystem, skills/plugins, hooksRepo-task automation, code-genWraps multiple frontier models, deep editor integrationBest-fit SecOps workInvestigation, enrichment, reporting, building team skillsBuilding automations / SOAR-as-codeMaintaining detection-as-code and tooling reposThe honest caveatAs good as the context and integrations you give itCode-centric; less optimal for team-wide implementationEditor-bound; built for engineers, not analysts

The practical takeaway: these platforms are broadly comparable, and you can do good work with any of them. Claude is a bit ahead today on collaboration and teamwork with shared skills, plugins, and MCP connections that let a whole SecOps team work from one toolkit. This matters in security operations, but the underlying mechanisms exist on every platform, just with more or less friction. And the space is moving fast enough that any assessment like this can change month over month. So invest in what makes sense for your org, including financially: pick one as your through-line, keep licenses monthly where you can, and stay flexible enough to switch as the space moves.

7. How to adopt: a phased playbook

Adoption is a three-to-six-month arc, not a switch you flip. The sequence below turns the stack model into a rollout that earns proof early and compounds from there.

7.1 Start with a task audit

Before you buy or build anything, map where AI actually helps. Survey the team's daily work and sort each task on two axes.

First, build or adopt: is this business-specific work you'd build a custom skill for, or generic work an existing tool already handles? Second, eliminate, accelerate, or keep: do you want it off your plate (autonomous), done faster with you in the loop (copilot), or left exactly as it is?

The audit template (downloadable). One row per recurring task:

Task nameTime (% of day)Impact (L/M/H)How it's done todaye.g. Exec summary for an incident report10%HPull the case timeline, expand affected users/devices/IPs, chase stakeholders for context, format for the board…

That last column pulls double duty. It's how an analyst describes the task today, and it's the spec you'll later hand Claude to build the automation, or for the work you want to take off the team entirely, the brief you hand your AI SOC platform to run autonomously.

7.2 Roll out: pilot, then champions, then scale

Start the pilot where the audit shows the biggest time sink with the clearest payoff, which is usually triage and enrichment or reporting. Prove the value on one workflow before you widen it.

Give one champion access first. They build the shared skills and connections, and access opens up gradually after that, so team members can contribute through PRs into the shared setup. Pick a builder-teacher too: one person who both develops the shared skills and prompts and trains the team on using them, or two people splitting those roles with a clean handoff.

Treat enablement like phishing-awareness training. Short, practical sessions that get everyone to a working baseline. Leaving a tool on people's desks is not a rollout.

7.3 Build the shared toolkit

The difference between siloed individual wins (stage 1 on the maturity curve in Section 3) and a compounding team (stage 4) is a shared toolkit.

Build a team plugin to standardize the skills, scripts, and connections, so people collaborate instead of each reinventing their own prompts. Ship a few starter skills; reporting is a natural first one, and the library grows as the audit surfaces more repeatable tasks. For connecting AI to your stack, a few tactics earn their keep. Prefer a normalized layer over wiring AI into a dozen tools one at a time, since it abstracts the per-tool query languages like KQL, SPL, XQL, and SDL. (Raw tool access isn't investigation; the AI SOC layer is what gives you the normalized, forensic version.) Where a system has an API but no MCP, write a script. Where it has no API at all, drive it with a browser MCP like Claude for Chrome.

For your builder-teacher. Write outcomes, not steps: give the goal and the tools and let the agent find the path. Don't pin MCP tool names; reference the capability so skills survive upstream changes. Keep your conventions in CLAUDE.md so quality holds as the team grows. Test honestly, which means unit-testing the deterministic parts like parsers and scripts, and hand-validating the probabilistic agent steps. Push noisy work to subagents to protect context. And guard the risky skills with disable-model-invocation: true for anything that has side effects. And keep the plugin in a private GitHub repo: it makes shipping it internally painless, and every update syncs to the whole team. Anthropic's guide to building plugins covers the mechanics.
The minimum guardrails. Secure adoption is its own document (Section 1 points you to NIST and OWASP), but don't start the pilot without five basics: the AI acts with each analyst's own identity and scoped credentials, never a shared super-user account; every AI action gets logged; anything with side effects — containment, ticket updates, outbound messages — goes through an approval gate; a human signs off on every consequential verdict; and unknown samples only ever detonate in a contained environment.

7.4 The first 90 days

A 30/60/90 plan, with your starting point keyed to your Section 3 maturity stage:

Adoption playbook infographic showing the first 90 days in three phases: prove value, connect and widen, compound

Where you start depends on your Section 3 stage: a Stage 1 team spends day one on the audit and a first skill, while a Stage 3 team that already runs an autonomous layer can jump straight to closing the tuning loop.

Three ways this goes wrong. First, skipping the middle layer: wire AI straight into detection tools and the agent can't sustain 24/7 investigation, so you get a disappointing demo and write off the whole idea. Second, ignoring the economics: per-alert investigation cost balloons, the team quietly deprioritizes low-severity alerts, and the coverage gap is right back. Third, refusing to adopt a platform at all: trying to out-build the frontier-model companies instead of giving Claude or Codex a solid foundation to run on.

8. The human side: what changes for your team

Adoption fails on people more often than on technology, so plan for the human shift on purpose. Two things change.

The analyst's role moves up the value chain, from doer (manually gathering, pivoting, assembling) to reviewer and supervisor of the AI that does the gathering. The skills that matter shift with it: less rote tool-jockeying, more judgment, more tuning, and a feel for when the machine is wrong. Say this out loud, and invest in the new skills. That's what the builder-teacher and the enablement sessions in Section 7 are for.

And take the "will AI replace me?" question head-on, because dodging it breeds the quiet resistance that stalls rollouts. The honest answer is augmentation. AI takes the drudgery, the 3 a.m. alert grind, the copy-paste enrichment, the report assembly, so analysts can do the higher-value work that drew them to security in the first place. Capacity-bound teams don't shrink. They finally cover the ground they never could.

From doer to supervisor: how an analyst's workday is re-spent when AI handles execution

9. Measuring it: the board narrative

If you can't measure it, you can't defend the spend. Baseline before you start, then track the movement.

MetricHow to baselineWhat good movement looks likeMTTR (mean time to respond)Current average across recent incidentsA material drop as triage and containment are automatedAnalyst hours reclaimedHours per week on triage, enrichment, reportingHours shifting from gathering to decidingAlert coverageShare of alerts that get a real investigation todayToward 100% once the autonomous layer is connected — the single biggest lever on breach risk, since the low- and medium-severity alerts a capacity-bound SOC skips are where intrusions hideBacklog burn-downDetection and phishing backlog sizeSteady reduction as the copilot loop runsCost per incidentFully-loaded analyst time per incidentDown as work shifts to the autonomous layer

Illustrative math for a 10-analyst SOC. Say 2,000 alerts a day and a team of ten. Human capacity gets a real look at maybe 200 of them, call it 50 analyst-hours a day, with 90% of alerts unexamined. While the autonomous layer investigates everything and escalates ~2%, the team reviews ~40 fully-worked cases a day which comes out to roughly 10 analyst-hours. That's ~40 hours a day handed back or about five analysts' worth of capacity redirected from triage to hunting and tuning, all while coverage goes from 10% to 100%. Run the same math with your own volumes; the shape survives even conservative inputs.
Alert coverage funnels comparing today's 10% human-handled coverage with 100% investigated by the AI SOC

Tokenomics

A word on tokenomics, since it's where ROI quietly leaks. The architectural point is that an AI SOC layer means you're not running every one of thousands of daily alerts through metered general-model tokens. The autonomous-layer vendor owns the problem of triaging at scale both accurately and cheaply, so your copilot tokens get spent only where humans actually engage, on the escalations. Wire your copilot to investigate every raw alert directly and the per-alert cost is exactly what pushes teams back into ignoring the quiet alerts. On the copilot side, tooling helps too: open-source proxies like Headroom and RTK compress tool outputs before they reach the model and cut token spend by well over half without hurting answer quality.

For proof points: mature autonomous triage can close on the order of 98% of alerts on its own, and because the layers compound, every human tuning decision cuts tomorrow's escalation load. One large US-based pharmaceutical company runs this model at a mean time to respond of 2.5 minutes. The economics get better over time instead of worse.

10. Common concerns

Answering the obvious objections plainly buys credibility. Dodging them spends it.

Will this replace my analysts? No. It changes what they do. AI handles the volume and the assembly; analysts own the judgment, the tuning, and the decisions that carry business context (see Section 8). Capacity-bound teams use it to finally cover ground they couldn't. And it rarely means shrinking a team dramatically; what changes is the hiring math — instead of throwing more bodies at rising alert volume, the team you already have absorbs the growth.

What about hallucinations or wrong answers in a security context? It's a real risk, and you manage it by design. Ground the AI in your actual tool data instead of its training-time "knowledge," keep a human on every consequential verdict, and require evidence with each claim. The AI SOC is a faster, better-documented “outsourced SOC”, not an unattended actor.

What data do these platforms actually see? Less than you might fear, if the architecture is right. The copilot works the escalated cases, not your raw alert firehose, and it acts with each analyst's own identity and credentials — it sees what that analyst is already cleared to see. The rest is standard procurement homework for whichever copilot vendor you pick: enterprise data controls, no-training commitments, SSO, audit logs. The secure-adoption specifics are out of scope here; Section 1 points to the right frameworks.

We tried AI tools before and they didn't deliver. Usually that's a tech stack problem (Section 3): a chatbot bolted on, or an agent wired straight into detection tools. The stack model, an autonomous layer for volume and a copilot for judgment, both grounded in your context, is what makes it stick. It's also a question of maturity: the category has moved fast, and in 2026 the best AI SOC technology can match or beat best-in-class MDR on speed and depth. A tool that underwhelmed a year or two ago isn't the tool you'd be buying today.

Won't the token bill explode? Only if you point metered general-model tokens at the alert firehose — that's an architecture mistake, not a law of nature. Keep the volume work on the AI SOC layer, whose per-endpoint economics don't punish alert count (Section 9), so copilot tokens are spent only where humans engage. Standardize the copilot side through the team plugin and skills, so model selection is deliberate — a fast, cheap model for mechanical steps, the frontier model where judgment matters — instead of every analyst defaulting to the most expensive option. Compression proxies like Headroom and RTK (Section 9) trim what's left, and monthly licenses (Section 6) let you rebalance as pricing moves.

How do I justify the ROI to the board? With the metrics in Section 9: MTTR, analyst hours reclaimed, coverage moving toward 100%, and cost per incident, baselined before and tracked after. Lead with coverage: moving it toward 100% is the metric that most directly cuts breach risk, because the low- and medium-severity alerts a capacity-bound SOC skips are exactly where real intrusions hide.

11. Where to go next

You don't have to boil the ocean. You have to climb one rung. Run the task audit, pick the highest-payoff workflow, and stand up the copilot layer with one shared skill. That's a 30-day move.

If you only do five things:

  1. Run the task audit — one week, the whole team, one row per recurring task.
  2. Stand up one copilot with one shared skill aimed at your biggest time sink.
  3. Baseline MTTR, coverage, and analyst hours now, before anything changes.
  4. When you add the autonomous layer, demand full coverage — every alert, at economics that don't punish volume.
  5. Name the champion and the builder-teacher, and train the team like you run phishing awareness.

When you do evaluate an AI SOC platform, these are the questions worth asking:

  • Does it investigate every alert, including low and medium severity, or only the loud ones?
  • How deep does the investigation go, real forensic evidence and a reasoned verdict, or just data access?
  • Who owns the knowledge layer, your detection logic and case history?
  • How does it hand off to your copilot, with the full trail, or a black-box "look at this"?
  • How do the economics hold up at scale (per-alert cost)?
  • Is it built for a human in the loop, or for automation theater?

Further reading

If you want to watch the autonomous layer and the copilot handoff work together on your own alerts, contact our team for a demo.

Itai Tevet

Co-founder and CEO of Intezer, Itai is on a mission to revolutionize how SOC teams investigate and respond to cybersecurity incidents. He previously led the cyber incident response team for one of the world's most targeted organizations. Itai combines his expertise in AI and security to advise security leaders at Fortune 500 companies on how to defend against threat actors in the AI era.

In this article

Share article

Related Articles

Company News

4 min

Intezer Workflows. The AI SOC is now complete

Detect, triage, investigate, respond. The entire SOC lifecycle now runs in one platform with AI executing and humans supervising. 

CISO

CISO Playbook: Putting Claude to work in security operations

This playbook is for security leaders who know AI belongs in the SOC, but need a practical model for where it actually fits. It’s written for CISOs, SOC leaders, detection engineers, and security teams dealing with alert volume, manual triage, reporting drag, and pressure to justify AI investment.

Company News

5 minutes

Loop engineering comes to the SOC: Introducing the Intezer Org Brain

Organizational context in an AI SOC is table stakes. Org Brain is very different. It learns, it recalls, it fetches what it's missing, and it gets sharper with every alert it touches, all autonomously.