0:00
/

Cybersecurity is entropic too!

A discussion with Surf AI CTO Roie Cohen Duwek

Every system is an imperfect model of messy reality -- the messier the domain, the more the system costs to build and the more it frustrates its users. [1] GenAI lowers the cost of absorbing the mess: it reads the PDF, reconciles the schemas and constructs the rule. Projects that never cleared a business case now might -- so look for entropy.

Cybersecurity is entropic. [2]

Entropy concentrates where judgment resists reduction, in cyber risk assessment, incident response and forensics, and crisis management, and the least in the plumbing: authentication, identity lifecycle, log management, key management, facility access.

Initiatives a company might pursue reconcile, correlate, consolidate or unify records that disagree: continuity plans against recovery evidence, asset inventories against cloud findings, telemetry against case files.

Today we interview Roie Cohen Duwek, co-founder and CTO of Surf AI, which focuses on exposure management. Which vulnerabilities matter, who owns the fix, what breaks when someone applies it. He echoed the model more than I expected.

The model’s first lever for vulnerability management is to assign an owner to each finding, and its second is to assemble scanner output, asset records and compensating-control documents into one package. Roie described the same problem from the other direction. The hard question, task ownership, asks who has to take part in a given fix, and the answer starts with knowing whether an identity is an employee, a contractor or a service. That is heterogeneity, and his fix matches the model’s: normalize roles from Entra, Salesforce and AWS into one entity, and tie every account to a person. He calls it subject defragmentation. [6]

The model scores vulnerability management as low in rule complexity and medium in data opacity. Roie describes a domain with more findings than any team can fix. The model scores the work as organizations do it today: scanners emit findings, and patch rules rarely change. Roie points at what a fix touches: attack paths, compensating controls, owners, ripple effects. The entropy sits in the connections, not the rules, and the model may under-weight it.

Agentic Exposure Management


In this section:

  • What Surf AI does, and why it focuses on exposure management rather than incident response

  • Why vulnerability volume and a changing attacker profile overwhelm manual approaches

  • How agents combine context, attack paths and ripple effects to prioritize and remediate

  • Where humans remain accountable, and how trust in the system grows


James Kaplan: Hi there. This is James Kaplan with another ProsaicTimes video podcast. I’m here with Roie Cohen Duwek from Surf AI. I hope I pronounced that okay. Roie, why don’t you introduce yourself and tell us a little bit about Surf AI?

Roie Cohen Duwek: Happy to be with you, James. My name is Roie Cohen Duwek. I am a co-founder and CTO at Surf AI, where we help large organizations operationalize security events, especially non-SOC operations, everything that has to do with exposure management, in a context-aware system of agents.

James: Let me see if I can summarize this briefly. You’re trying to bring agentic execution to the security operations center. Is that a fair description?

Roie: It is, but I would say that we look mostly at everything that has to do with exposure management. That includes exposure related to IAM, vulnerability management, or really any type of exposure management. It’s not the incident control needed once penetration has been achieved, but mostly what comes before that and during hygiene cycles.

The bulk of the value we bring to clients is the automation and operation of security hygiene. It’s achieved by making context available to the system of agents that run the ship.

James: Let me see if I understand the problem you’re trying to solve. There’s no end of tools that will identify vulnerabilities, and most organizations I know of are swimming in them. They have a hard time understanding which vulnerabilities are more important than others, from both a business and a technical standpoint. They have a hard time understanding how to fix them, and a hard time executing on the fixes. Every institution I know of has more vulnerabilities than it can possibly fix. Is that a fair description of the problem you’re trying to solve?

Roie: That’s part of it, for sure. Vulnerabilities are becoming a commodity. Recent stats show vulnerabilities surpassing credential abuse for the first time in twenty-plus years as the leading vector [7], and more vulnerabilities are being used as zero days than as one days in new exploits. [8] AI is also being used increasingly by smaller groups, not limited to state actors anymore, which means the entire landscape is changing. We see a new attacker profile, which we can discuss at length if you would like.

James: Put the landscape aside for a minute. Explain a little bit how this works. Why does interrogating and executing on vulnerabilities work agentically? Why is that a good use case for an agent, and what is some of the logic the agent employs?

Roie: Are you asking why it works, or why we need it in the first place?

James: Why does this work? If you were a CISO, why would you trust this type of approach?

Roie: I would say this type of approach is a reaction to the changing landscape, in which vulnerabilities are much more accessible.

James: We get the changing landscape. My question is, how does this work?

Roie: First, we need to understand the client’s environment very thoroughly: where we are vulnerable and where we are not, and which attack paths are present and realizable in that environment. That understanding affects the priority of the findings we have from different sources. We need to understand what compensating controls can cut those attack paths short, and the ripple effects that fixing those vulnerabilities and other exposures may cause. We need to discern the ROI and understand the context. Then we lead assisted remediation through workflows with the users, and cyclically reevaluate health.

That is a very complex thing to do by hand, and done manually it is naturally limited to a very small number of such findings at a time. As the number of findings rapidly accelerates, we need a reaction that can respond in real time. It needs context that would otherwise be available only to expert humans, and it needs to work relentlessly, keeping the organization safe.

James: What gives you confidence? How do you think about which cases an agent can triage or develop a remediation plan for itself, and when it needs to escalate to a human?


One of the most important parts is understanding the ripple effect. It’s important to identify ahead of time, because it means we can anticipate not just what can happen if a bad actor reaches the asset, but also what bad things could arise if we do the right thing and patch the vulnerability.


Roie: When we have an exposure we want to remediate, we ask ourselves the following. First, we need to understand exactly what we’re talking about, so we need to identify and classify the assets involved. We need to understand how this organization specifically handles remediations of this type. We need to know who needs to be involved: which humans take part, and which stakeholders need to make decisions. Then we assist the users in remediating the exposure.

One of the most important parts is understanding the ripple effect. It’s important to identify ahead of time, because it means we can anticipate not just what can happen if a bad actor reaches the asset, but also what bad things could arise if we do the right thing and patch the vulnerability. Then we can plan for them.

Most clients today want a human involved in every operation initially, until they gain trust in the system. Where they want a human to remain involved is mainly where there is a high risk of system disruption or privacy breach, anything where they want real accountability. In those cases they will employ humans as decision makers.

Agents and the Security Analyst Role


In this section:

  • How the analyst’s job changes when agents handle analysis and prioritization

  • The two goals of agentic defense: shrinking the attack surface and reducing blast radius

  • Why SOC headcount is unlikely to fall significantly

  • How span of control differs between incident management and exposure management


James: How does the job of the cybersecurity analyst or operations center technician change as you move to a more agentic security organization?


An attacker is going to get into your systems eventually, and once that happens, you want to catch them fast and stop them from causing too much damage while they’re there.


Roie: I would say it changes dramatically. First of all, analysts can do much more with their time than they used to. They can employ agents to analyze attack paths, and to identify and prioritize various exposures. They can use various systems and agents to deal with penetrations once those occur. They can also use agentic systems to achieve two goals. One is to minimize the attack surface of the organization. The second is to reduce the blast radius, because in today’s reality you have to assume eventual penetration. An attacker is going to get into your systems eventually, and once that happens, you want to catch them fast and stop them from causing too much damage while they’re there.

James: How does the skill model change for cybersecurity analysts working at large enterprises? Does it mean you need people who spend more time on the more complicated problems? Does it mean you need someone managing fleets of agents? How does the skill set evolve, in your thinking?

Roie: We already see a change happening in security operations center teams. It affects the number of people hired, because strong security analysts can today achieve much more using agents. But in my opinion, we’re not going to see a significant reduction in force in the SOC. On one hand, the analyst can do much more with agents than before. On the other hand, the attackers can also do much more than before.

We see a rapid increase not only in the speed and sophistication of new attacks, but also in the ability to chain different attacks together and to act relentlessly. Agents can act relentlessly over extended periods of time and encompass every security skill imaginable. In the past, a team of bad actor hackers might have one expert on the network and another expert on software vulnerabilities. Now every one of the agents can be an expert on all of those, and you can have many copies of each acting together over a prolonged period of time.

James: How do you think about span of control? Most people think of span of control as cybersecurity analysts per number of applications or parts of the business. Do you think about it as an analyst being able to manage a certain number of agents? How does productivity increase, and what’s your experience there?

Roie: Here I would divide the world into exposure management and incident management. In incident management, you would have several different agents collecting information and assisting you in analyzing what’s being hit, how to stop it, and how to isolate the penetration.

In exposure management, you’re going to have more and more agents managing workflows and sophisticated identification schemas. For instance, an agent identifies that an attack path exists within the system, decides where to stop it, contacts the people, and suggests the highest-ROI actions to make the attack path unviable. It can even implement those by itself, submitting the work for human verification and approval.

Mythos and the Shifting Threat Landscape


In this section:

  • Whether Mythos and open-weight models have already changed the threat landscape

  • The Hugging Face incident as an example of autonomous, multi-step agent attacks

  • James’s counterview that guardrails and open-weight lag mean the other shoe hasn’t dropped

  • The economics of attack: parallel, continuous work by ever-smaller groups


James: How does this all change in a Mythos-slash-Fable world? [9] Especially as more sophisticated open-weight models become available, attackers may be able to discover vulnerabilities more easily and create fairly sophisticated multi-step attacks, and there may be a greater velocity of vendor patches that people need to assess and address. A lot of people have hypothesized that this will require more agentic technology in the security organization. Could you explain how that works and how to think about it?

Roie: First of all, Mythos changed everything.

James: Let me challenge that. Has it changed everything yet?

Roie: I think it has, in a sense. Not Mythos itself, but what Mythos introduced: a long-horizon scale of models, which of course extends beyond itself. They are now available, and will be available in stronger and stronger varieties, to bad actors as well as to defenders. You can see that in the famous Hugging Face incident. [10] The agents were able to initiate the attack, collaborate and communicate over open channels, and relentlessly look for ways to exit their isolation. They acted, of course, with no morals. They used numerous zero days, discovered the attack paths on their own, abused them, and mapped the organizations to get the result they were after. All of that was casual. None of it was directed. No human told them to attack.

James: Let me offer the alternative view that Mythos hasn’t changed everything yet. The products from the frontier labs have significant guardrails around them. Maybe those can be jailbroken, maybe not, but that takes time, effort and skill. The open-weight models are some amount of time behind. Some people tell you it’s a month, some six months, some a little bit longer. But experientially, my impression is that the other shoe hasn’t quite dropped yet. Maybe at the next round of upgrades for the open-weight models, things get much more worrisome. I don’t want to counsel a lack of seriousness on anyone’s part. I just want to be realistic about what may be coming at us in the not-too-distant future.


Work that used to require a skilled team over a prolonged period can increasingly be done continuously and in parallel by agents.


Roie: I agree to an extent. As you said, the important change is not some magical new attack technique. It’s the economics. Work that used to require a skilled team over a prolonged period can increasingly be done continuously and in parallel by agents. Smaller and smaller entities are now using agents for attack. It used to be only nation states, then it moved to large attack groups, and now we’re seeing it even in small and medium-sized attack groups. Even in the Hugging Face incident itself, Hugging Face used an open-weight model to defend itself.

Measuring Success and Context Graphs


In this section:

  • How to measure agent effectiveness: work absorbed, cost, and quality of decisions

  • Why attack-path reduction is a better security metric than raw patch counts

  • How Surf AI’s context graph and specialized model enable deeper answers than text or tables

  • How the ontology stays constant across customers, and how hard it is to hydrate the graph


James: Right, because in this particular case the guardrails on the non-open-weight model were interfering. Let me ask this question: how do you think about agent effectiveness? Is it mean time to restore? Exception rate? Failure rate? How do you think about measuring the effectiveness of an agent in the cybersecurity domain? And how do you think about span of control? How many agents can someone manage?

Roie: I would say that agents are becoming better and better at decision making and at communicating with humans.

James: We all get that. But what would you say is true now?


So the right measurement is not only how much more we can do, but also how much more secure we are.


Roie: The effectiveness of an agent depends mostly on the breadth and depth of context available to it, and the expertise you give to each one. If you have an agent that is an expert, meaning it has the relevant information, focus and context in a certain field, say cloud security, talking to humans, ticketing systems or exploits in SaaS, and you also give it the relevant context in the organization, we find that it becomes very effective.

At the end of the day, I tend to look at agent effectiveness in terms of how much work agents take off the hands of the human operators who send them to do something. Say we used to be able to do a certain number of vulnerability patches a month, and now, without hiring anybody new, we’re able to do seven times that. That’s the measurement, and that’s what we wanted to achieve. Of course we have to look at how much we spend on those agents, but usually that’s not very hard. We would gain seven times the workflow, and importantly we could also be making much more calculated and better decisions on where to apply those patches and how to invest our effort. So the right measurement is not only how much more we can do, but also how much more secure we are.

James: How do you measure how much more protected you are? Is it pace of vulnerability remediation? Percentage of open vulnerabilities? How do you think about that?

Roie: In my view, looking solely at the number of vulnerabilities we patch, or the number of critical or high vulnerabilities, can be misleading. A medium vulnerability that sits in an attack path leading to a critical asset can be more urgent to patch than a legacy critical vulnerability sitting out front that doesn’t lead to any critical asset.

So if we adjust priority to everything we can now look at, such as attack paths, and vulnerabilities that used to be impractical but are now usable by agents, then we could look at the number of vulnerabilities fixed. But I would look at the amount of attack surface we have open. I measure attack surface by the attack paths leading into the organization. Say we had 74, divided into critical, high, medium and low attack paths, and we disrupted enough of them to shrink our attack surface significantly. That can be a very strong measure of our success.

James: You mentioned a graph that you use for context. Tell us a little about the shape of this graph, what type of information it contains, and how you assemble it.

Roie: The thing I find most helpful for agents to do their job correctly is the right context at the right time. What we call our context graph, or context fabric [11], works like this: we ingest the relevant information from basically any business software, normalize it into a context graph, and map all the relationships between all the assets across that graph. Those connections can be static or temporal, and you end up with a lot of nodes and arcs that describe the organization.

Then we take all of that graph and create a specialized AI model analogous to an LLM. With an LLM, you take the elements of a language and create a model that allows you to interact with it. We do the same with the graph.

James: This gets to one of my core hypotheses: you get a much better answer reasoning over a graph than over text. The edges between nodes contain a tremendous amount of insight, and reasoning over them reduces the chances of problems in the conclusions.

Roie: Right. Deep answers include things like how this organization typically solves a problem, process harvesting, or identifying the possible ripple effect based on years of accumulated data. They include identifying the type of an asset, or the owner of a specific task, meaning the person who needs to be involved. If you do this naively over tables, or even a graph database, it can be very complicated, and you would have to create a very specialized model for every such question. With the model we created here at Surf, you can ask the graph the question. You can interact with the graph in a similar way to how you would interact with an LLM, and get to very deep answers very quickly.

James: What’s an example of an answer you can get there that you think is really interesting?

Roie: I’ll give two examples. The first is asset categorization. You want to understand whether a specific identity is human or non-human, and what subtype. If it’s human, is it a guest, an employee or a contractor? If it’s non-human, is it an AI agent, a service, an application, a device? We created an ML-based model to solve this problem. It took quite a bit of time, over a month of work, and we created over sixty different submodels, some ML-based, some semantic, some statistical, some static heuristics. We got to an accuracy of over 98%.

Then we started anew using this model, without giving it any of the features that were available to the one created before, and in half a day we got 96% accuracy. [12] That’s a quantum leap in the ability to create new answers to deep questions.

Another example is task ownership. It is significantly more complex than asset ownership. It’s not just asking who is responsible for an asset. It’s asking who needs to be involved in this specific operation on that specific asset, and the answer can be more than one person. The first example could be done traditionally, but it would take a lot of effort. For task ownership, because of the variety of tasks, assets and organizational structures involved, I can’t even imagine a way to do it without such a model.

James: How much does the ontology stay the same across different companies, and how much does it need to be tailored or customized?

Roie: Our ontology stays the same across customers. We don’t change it from customer to customer. It is built to hold information from any business application, and it’s very extensible. The goal is that our models can run on it, agnostic of the platforms beneath it. Roles, for instance, are very different in Microsoft Entra, Salesforce and AWS, but they all get normalized into the system with their different qualities and limitations. When the models run on them, they see one same entity across all these dimensions, called the role.

James: How hard is it to hydrate the graph? To what extent do companies have this data readily available, and to what extent does it take a lot of work?

Roie: Companies do have the data. It’s interesting that in many cases we come back to companies with insights and findings on their own data that they were completely unaware of. We showed one company that it had thousands of unapproved applications running in its system, and they were completely unaware of it. [13] So companies usually have the data, but they are not always aware of the data they have.

James: So it’s a data correlation problem as much as anything else.

Roie: One of the data correlation problems required for operation is associating assets with humans. You want to say that all of these identities belong to one James Kaplan, all of those identities have these accounts, and these accounts work on these assets, and to do that across all the different platforms and systems you use. We call that subject defragmentation. It can become very complex in certain organizations, so data correlation is a very big part of it.

James: Any concluding thoughts, or anything really important we didn’t touch on that you think people should understand?

Roie: A concluding thought would be that the security landscape is changing in a way that makes attacks much more economic, relentless and fast, and conducted by automated processes. To be secure, defenders need to think the same way. What can we automate? How can we act relentlessly, cyclically and continuously to ensure the attack surface is as minimal as we can get it, and that if a bad actor does get in, the damage done or the data stolen is minimal?

James: Terrific. Thank you so much.

Roie: Thank you, James.

What should a CIO or CISO do with this?

Figure out where the entropy exists in your model, what leading or lagging indicators you want to address and what GenAI deployment patterns you might use. Score one of your own processes on the eight markers before you fund anything. Vulnerability management scores high on heterogeneity and modest on rules. Its value comes from the reconciled record, and the agent only reads what the record holds.

Measure paths, not patches. Count the attack paths into the organization, not the findings closed. Findings closed is a weak signal: Verizon found that organizations fully remediated only 26 percent of known exploited vulnerabilities in 2025, down from 38 percent, and the median fix took 43 days. [14] Closing more tickets does not shrink the surface.

Decide now what an agent may do alone. Start with a human in every operation, and the model rates run-time execution the riskiest pattern. Write the list before the pilot does it for you: what an agent does without approval, what needs approval, and who owns the exceptions. A materiality call stays with a person.

Footnotes

[1]: By entropy I mean the gap between the messy information a business runs on and the formal systems built to model it: data that is ambiguous, missing, unstructured or scattered across systems, and rules that depend on judgment, vary by case or depend on one another. The more of it a domain has, the more it costs to automate.

[2]: I ran the entropy model on it. It scores 41 sub-domains, from third-party risk to physical access, on eight markers: four of data opacity (ambiguity, incompleteness, lack of structure, heterogeneity) and four of rule and path complexity (discretion, rule volume and velocity, path variance, rule interdependency). Every rating cites a source, and an adversarial reviewer challenges each one. Then it designs an AI opportunity for every instance of entropy and estimates the hours freed.

[3]: ISO 22301:2019, Clause 8.4, requires documented business continuity plans and procedures. NIST SP 800-34 Rev. 1, Section 4, covers the information system contingency plan.

[4]: Form 8-K Item 1.05, added by the SEC’s Cybersecurity Risk Management, Strategy, Governance, and Incident Disclosure rule (adopted July 2023), requires disclosure within four business days of the company determining that a cybersecurity incident is material.

[5]: Article 23 of Directive (EU) 2022/2555 (NIS2) requires an early warning within 24 hours of an entity becoming aware of a significant incident.

[6]: Roie’s term for matching the identities, accounts and assets that belong to one person across all of a company’s systems into a single record; who owns what has one answer.

[7]: Verizon’s 2026 Data Breach Investigations Report puts vulnerability exploitation at 31 percent of initial access, up from 20 percent a year earlier, against 13 percent for credential abuse. Verizon describes it as the first time in the report’s 19 years that exploitation has led. Roie says twenty-plus years; the DBIR’s count is 19.

[8]: The zero-day versus n-day comparison is Roie’s spoken claim. I found no published count that supports it. The nearest figures: Mandiant’s M-Trends 2026 puts the mean time to exploit at an estimated minus seven days: exploitation began on average a week before a patch existed. Google’s threat intelligence group counted 90 zero-days exploited in the wild in 2025, 48 percent of them in enterprise technology (Look What You Made Us Patch).

[9]: I have written about Mythos twice. Is Mythos the Sputnik moment for AI in enterprise technology? (June 21) argues that remediation, not discovery, is the bottleneck. Trading bad inefficiency for good inefficiency at the Technology Leadership Forum (May 10) concludes that “Mythos is cause for determination, not panic.”

[10]: The account matches the primary sources. Hugging Face’s July 16 disclosure and technical timeline describe an intrusion by “an autonomous AI agent driven by a combination of OpenAI models,” about 17,600 recovered attacker actions between July 9 and 13, and no human directing the individual steps. Hugging Face believes the agent was trying to cheat an evaluation by stealing the test solutions. OpenAI’s account and the METR and Redwood Research investigation cover the same events from the other side. Hugging Face used an open-weights model, GLM-5.2, to decipher most of the agent’s payloads, which is Roie’s point about defense a few turns later.

[11]: Surf’s name for its context graph: information ingested from business software, normalized, and linked into one graph of assets and the static or time-varying relationships among them. Roie describes training a specialized model on that graph, the way an LLM is trained on language; agents query it. For the general idea, see Context graphs raise the waterline for machine judgement.

[12]: The 98 and 96 percent accuracy figures are Roie’s own, from Surf’s work on identity classification. I could not find a published test set, baseline or methodology for either. Read them as his report of internal results, not as benchmarks. The comparison he draws is the cost: more than 60 submodels over a month for the first model, half a day for the second.

[13]: The “thousands of unapproved applications” is Roie’s account of one unnamed customer. It cannot be checked independently, and I have left it attributed to him.

[14]: The figures are from Verizon’s 2026 Data Breach Investigations Report: 26 percent of vulnerabilities in the CISA Known Exploited Vulnerabilities catalog fully remediated in 2025, against 38 percent the year before, and a median of 43 days to full remediation, up from 32.

Discussion about this video

User's avatar

Ready for more?