Why Gateways Break AI Operations (And How to Actually Secure Agents)

View Show Notes and Transcript

When an AI agent deviates from its mission, it doesn't trigger a traditional escalation of privilege or alert your SIEM. It fails silently, behaving exactly like a normal user and looking at system logs won’t tell you if its intent was malicious or simply misguided.  In this episode, Ashish sits down with Hanah Darley, Chief AI Officer and Co-Founder at Geordie AI, to discuss the reality of securing an agentic enterprise. Drawing from her background in government intelligence and psychology, Hanah explains how cognitive biases cause security professionals to mistakenly apply legacy cybersecurity frameworks to non-deterministic AI. We unpack why routing agent traffic through traditional gateways introduces crippling latency, and why organizations must focus on "context engineering" to nudge agent behavior in real-time.  Hanah also challenges the "Human-in-the-Loop" fallback, arguing that in many objective research tests, humans actually produce worse outcomes than the AI itself. Finally, she outlines a pragmatic maturity model for AI security, starting with specific, contextual red lines and advancing to fully autonomous decision-making loops

Questions asked:
‍00:00 Introduction to Securing Agentic Ecosystems
‍01:50 Hanah Darley’s Background: From Government Intelligence to Chief AI Officer
‍02:50 Why "AI Security" is a Nonsensical Term
‍06:30 Defining "AI Native" vs. "AI First" Companies
08:00
The Problem with AI Gateways and Operational Latency
11:50
Why Logs Are No Longer Enough to Secure AI Agents
14:30
Understanding the Full Agent Lifecycle (Pre-Prod to Live Operations)
17:30
The Myth of the Security "Kill Switch" for AI
‍19:30 Why Agents Break Out of Sandboxes (Goal Pursuit)
‍22:30 Why System Prompts and Guardrails Aren't Foolproof
‍23:50 What is Context Engineering?
‍28:30 Psychology & Cognitive Biases in Threat Hunting
‍33:30 The Danger of Over-Relying on "Human-in-the-Loop"
‍38:50 Managing Multi-Dimensional Risk Across the Enterprise
‍43:30 A Maturity Model for AI Security Programs
‍47:30 Geordie AI’s Agent-Out Deployment Approach
‍51:00 Why Hanah Hasn’t Taken the AGI Pill

Hanah Darley: [00:00:00] We just default to like human is better. And in a lot of objective research testing, human is worse. This is foolproof. We have a container

Ashish Rajan: He can't get out of the room, but then obviously he broke out of the room

Hanah Darley: You're not the only actor in the room. It's not just the user. The agent has its own persona

Ashish Rajan: Maybe that's why people don't see ROI.

Ashish Rajan: We're still trying to apply AI to the old model because I don't want Bob, who's my best friend, to leave his job

Hanah Darley: some cases, malicious agent behavior looks like normal user behavior. I'm not taking the AGI pill. More AI is worse, and in some cases, less AI is worse

Ashish Rajan: Do you know where all your agents are? Do you even understand how do you apply security to an ecosystem that you don't understand?

Ashish Rajan: I just realized that most of us security people do not understand what a good scenario is. I had a conversation with Hannah Marie Daly. She's the co-founder and chief AI officer at a company called Geordie AI, and we were talking about how do you measure the behavior of an agent as to, is it doing the right thing?

Ashish Rajan: Do you know where all your agents are? Do you even understand how do you apply security to an ecosystem that you don't understand? We also spoke about how do you apply that to a security program that has existed for a long time, and some of the [00:01:00] biases we may have as a cybersecurity professional with decades behind us trying to tackle a new problem through pattern recognition, which probably doesn't work today.

Ashish Rajan: All that and a lot more in this episode with Hannah Marie, and I hope you enjoy this conversation of AI Security Podcast. And if you have been enjoying the episodes of the podcast for a while and you are here for a second or third time, I really appreciate if you drop the follow subscribe button, whichever platform you listen this on.

Ashish Rajan: We are on all podcast platforms like YouTube, LinkedIn, Spotify, Apple, or wherever you consume your podcast from. I hope you enjoy this episode, and I'll talk to you soon Hello, and welcome to another episode of the podcast. I'm here with Hannah. Thanks for coming in, Hannah.

Hanah Darley: Any time.

Ashish Rajan: Uh, I mean, I'm looking forward to this conversation also because I think there's some of the things we're gonna talk about has been on top of mind for me.

Ashish Rajan: But maybe to kick things off- Yeah ... if you could share a bit about yourself, your professional background.

Hanah Darley: Yeah. I'd love to. So my name is Hannah Darley. I'm the chief AI officer and one of the co-founders at Geordie, and my background is varied for someone in cybersecurity. So I started in government intelligence, which I did for about 10 years.

Hanah Darley: If my co-founder Henry would hear, he would [00:02:00] give you a ton of details that he should not share about that. Uh, but I just leave it at government intelligence. But lots of data, lots of analytics. And then I moved into cybersecurity and AI full-time for about five years. And a huge amount of that was working with enterprises on how to choose the right AI for the right problem, and also thinking about it from the security landscape and how do we think about cybersecurity in the context of an AI era uh, which has obviously become even more complex.

Hanah Darley: And so all of that feeds into my role at Geordie. We have been going now for quite a while. We're post Series A, and so I am having the lovely job of essentially making sure that all of our customers are able to enable agentic operations safely and securely.

Ashish Rajan: So this is fascinating, but also because we were all at BlackHat recently.

Ashish Rajan: Yeah. And one of the things that I noticed over there is the definition of what AI security's supposed to be is a bit murky. I think that's the word I'd go for it. I think everyone understood there is security for AI, AI for security, but the true definition of what is it that I can [00:03:00] build, what is it that the, a product company would build, what is it that is a gap today?

Ashish Rajan: Yeah. How do you define AI security?

Hanah Darley: So I actually think AI security is a completely nonsensical and unhelpful term Because there is so many different types of AI. There are so many different applications of AI when it comes to machine learning, when it comes to some of the deep learning applications, when it comes to even large language models, which we're all now obsessed with.

Hanah Darley: They all take such different principles and practices to make them safe and also to evaluate whether or not they're effective. And so if you said to me, "Oh, I'm an AI security company," I would say, "Okay, so you have to cover the entire spectrum of

Ashish Rajan: AI- Mm ...

Hanah Darley: as a business," which to me is like you have to have some, like- big business money behind you in order to

Ashish Rajan: be that- You're Anthropic at that, you're Anthropic at that

Hanah Darley: point in time Y- you're bigger than Anthropic.

Hanah Darley: Like- Oh. You're bigger than Anthropic. So I just think it's like a crazy term, because actually what people are usually talking about is one of two things. They're either saying, "I'm AI native as a product, and I operate in an AI-driven world-

...

Hanah Darley: Through [00:04:00] secure operations," or, or whatever I'm providing security for.

Hanah Darley: Yeah. Or I'm specifically trying to secure an aspect of AI, and within that aspect, I think you have to be definitive. I think you have to say, "I secure LLMs and how they operate. I secure prompts across multiple different streams. I secure algorithms across, machine learning inputs. I secure agents."

Ashish Rajan: Mm.

Hanah Darley: I think that's where you have to get to in order to start having a real conversation about what do you actually secure, and how does that work in practice.

Ashish Rajan: And I think that's a good distinction also because many people, especially CISOs as well, when they're building a security program- They're probably coming more from a, "I have an AI security program."

Ashish Rajan: Yeah. To, and to your point, like what is the AI security that you're trying to actually do?

Hanah Darley: Yeah.

Ashish Rajan: Right? And 'cause, and I'll probably take that example of yours and go one step further and say, what is the kind of AI you're consuming today that you should do, be doing security around- Yeah ... is what, technically what you're saying.

Ashish Rajan: If, as a-

Hanah Darley: Yeah ...

Ashish Rajan: person who's an end user, that's what they're think- they should be thinking about.

Hanah Darley: Yeah, 'cause also there's a piece of it that's [00:05:00] like I have AI in my detection response products now because that has become table stakes. Yeah. Everything that we do has to have AI in it from a detection response capability.

Hanah Darley: So then that's part of my AI security program, but it's not technically securing my AI in most cases. Mm. It's securing the rest of my estate. Yeah. And so, but it gets all looped into the same kind of bundle. And so I think we almost have to talk about, are we talking about securing the AI itself that we're using?

Ashish Rajan: Yeah.

Hanah Darley: Or are we talking about securing our operations with AI? And sometimes those have like a Venn diagram overlap, and sometimes they don't. Yeah. But I, I think that's one of the most helpful things for like security leaders. Yeah. Because that will also affect what you own, right? Like some of that from the, what we're securing for the estate, you definitely will own.

Hanah Darley: But some of it, like if you are not using your own large language models or your own in-house models-

Ashish Rajan: Yeah ...

Hanah Darley: you may need to worry less about securing LLMs. Mm-hmm. And you might need to worry about securing end user usage of specific harnesses, so like agents and things like that. Do

Ashish Rajan: you find that isn't an interesting, I guess, extension there that is [00:06:00] worthwhile calling out with the way agents operate and the way models are used?

Hanah Darley: Yeah.

Ashish Rajan: AI native companies, I can call myself an AI native company just because I use LLM. Yeah. And while it could just be I open chatgpt.com and go, "Hi, tell me what's the weather today," right? What do you- I'm AI first. Yeah. Actually, yeah, sorry. I'm AI first. So like that's actually a better one to go for.

Ashish Rajan: So-

Hanah Darley: Yeah ...

Ashish Rajan: I'm curious also in terms of how you see people who are truly AI native versus the AI first versus the AI bolt-on. How does me as a security practitioner or the CISO differentiate between the true AI native? 'Cause I guess my, my thinking here is that I would like preferably like to work with someone who's AI native.

Ashish Rajan: Yeah. And I think AI native is someone to what you said, everything about you is starting with the AI e- ecosystem first.

Hanah Darley: Yeah. I would define it as if you took the AI element away-

Ashish Rajan: Yeah ...

Hanah Darley: your product is useless. If that's not true, you're probably not AI native. And I would then say you probably are AI first.

Hanah Darley: So in our mechanisms and [00:07:00] toolkits, we will go first to an AI solution to help.

Hanah Darley: But that's not actually the premise of our product. That's not actually what we're building. We're not building AI, or we're not servicing AI. Right. We're helping do something else, and AI becomes a toolkit within that box.

Ashish Rajan: Would you say the current methodology or the way people think about this today, like, there is obviously there's a proxy people think of- Yeah ... there's a gateway people think of. Yeah. There is LLM firewall- Yeah ... people think of. There's like so many terminology. So could you just maybe highlight the approaches people have to, to what you said, and which one is the right way to go about, or should I look at all of them as a CISO?

Hanah Darley: So I'll narrow just into like the generative AI era, because you could talk about algorithm management.

Ashish Rajan: Oh, yeah, yeah. So I'm

Hanah Darley: narrowing

Ashish Rajan: it. You're very different to that. Yeah, yeah. I'm

Hanah Darley: narrowing it.

Ashish Rajan: Yeah, yeah. We're AI first, but we're

Hanah Darley: GenAI first. We're AI first, but we're really just talking about GenAI.

Ashish Rajan: Yeah, yeah. Literally. Yeah.

Hanah Darley: So when we're talking about generative AI, there tends to be one primary approach, which is if I can put some sort of a solution in between the user and the output- Yeah ... or the user and [00:08:00] the model-

Ashish Rajan: Yeah ...

Hanah Darley: then I can make it safer. Now, 9 times out of 10, that is going to be either an LLM gateway, an API gateway, or in newer age technologies, so where we're talking about agents, it might be like an MCP gateway.

Hanah Darley: So I'm specifically putting myself between the user and the tools- Yeah ... or the agent and the tools.

Ashish Rajan: Yeah. Um,

Hanah Darley: and in some cases you kind of try to do an amalgamation of all of them, and I just, I'm a network gateway, so I'm gonna be between everything. Yeah, yeah, yeah. Um, but the challenge with gateways is kind of a chicken and the egg.

Hanah Darley: So you have to know what you have in order to route it through. But predominantly in the generative AI space, the key challenge for most teams is uncertainty. Mm. I don't know everything that's happening in my estate, and to some degree that's true in other network operations, but it's really, really prevalent in generative AI because you have so many end users with different purposes and goals.

Ashish Rajan: Mm.

Hanah Darley: And so when you then say, "Okay, I have to know what I have in order to route it through a gateway," that's just- baseline challenge. The next challenge is scalability. And if everything has a gate, think about it like going through Heathrow.

Hanah Darley: If every five steps you had to do a [00:09:00] security check- Yeah ... and you had to have a gate, and they had to scan something or check something, you would take hours.

Hanah Darley: And I'm saying that, and just because you already take quite a long time to get through. But if you're thinking about it from how do I scale generative or agentic operations, one of the key challenges is latency and operational friction. Gateways introduce both. And so where gateways are super valuable is in high-stake scenarios where you would rather have an outage than an incident.

Hanah Darley: Oh, yeah. And in most businesses, it's the opposite. Yeah. I would rather have my CEO who can get online, who can use their agent, and then if I have to deal with a security incident, I'll deal with it.

Ashish Rajan: Yeah.

Hanah Darley: But when you're talking about how do I make agentic operations work at scale, most gateway solutions, in addition to that problem, also have a fail open and fail close choice.

Hanah Darley: Mm. Which means you just end up going back to a binary system for what is a very dynamic actor.

Ashish Rajan: But would you say, I love the analogy also because the whole outage versus, hey, am I, I think I can message. I don't wanna get a call from my CEO that he or she can't use my agent.

Hanah Darley: Can't do anything.

Ashish Rajan: [00:10:00] Yeah. But at least- That is a better s- that is probably the worst scenario.

Ashish Rajan: I would say is the approaches to these things, which is whether it's a gateway, a proxy probably maybe we can go a step further for companies that are very AI forward, like they have custom harnesses- Mm-hmm ... perhaps some kind of copilot- Yeah ... internal chat systems. Yeah. Like, people have obviously taken this to a long, uh, quite, quite a bit.

Hanah Darley: Yeah.

Ashish Rajan: What model makes sense? Am I using both approaches or does one have a blind spot over the other? Like what's, what's the recommended approach then?

Hanah Darley: Yeah. So I would say you need something that's more dynamic and more harness level when it comes to agents and their operations if you want to actually scale them.

Hanah Darley: Yeah. If you want to use agents in a really small corner of your business, actually a gateway model would probably work really well for you because it'd be really protected. Mm-hmm. Where I would say a gateway's more helpful is when you're thinking about LLM choices, because what we've seen a lot of is like LLM routing LLM preferences, and having a real focus on the ability to screen data from specific providers- Mm-hmm

Hanah Darley: and have data privacy agreements [00:11:00] with model providers. If that is your concern, a gateway at the LLM inference point is a great point to make. But as soon as you get into the agentic operations, most gateways just tend to slow your operations down to a point that the agent becomes ineffective. And I think one of the key challenges around this is it's not the same technology.

Hanah Darley: It ain't your nan's traditional software. Like it's something different, and so trying to treat it like it's just one point in time traffic, it's just one single mechanism that I need to approve and constantly push packets through, it just doesn't work that way. It's decision-making boots on the ground, and obviously it's predictive.

Hanah Darley: It's not sentient.

Hanah Darley: Yeah, yeah. But it is really different from hard-coded logic that's executed consistently over time like traditional software.

Ashish Rajan: Do you find that it's-- I love the analogy that we can't... obviously there is a-- I think what you're referring to also is your fundamental controls still work, but the problem that we're trying to solve here with the newer approach of using a gateway depends on whe- what is-- like are you okay to introduce latency?

Hanah Darley: Yeah.

Ashish Rajan: What is the, like I guess to what you said about the unknowns earlier, I was talking to [00:12:00] someone recently, and we were talking about this thing where, what am I looking for in the logs? Like, you know, to your point- Yeah ... I've figured out a way. Let's just say I've, whatever MCP proxy- I've got it. Yeah.

Ashish Rajan: I've figured out something, and now I have the OTel logs coming in. I'm like, "What am I looking for in this?" Do you even know? What do you recommend people-- Obviously, I know you guys have a solution for this, so I'm just curious in terms of- How do you guys approach this, and why have you gone down that path as well?

Hanah Darley: Yeah, I think one of the most important things that you're pointing out is logs aren't the answer. Yeah. And it's hard because in a lot of other security situations, logs are the answer. Yeah. Logs are self-explanatory. They're self-evident. They tell you exactly what's happening.

Ashish Rajan: Yeah. I'm not reading English at that point.

Hanah Darley: I got it. Like, I'm good. Yeah, yeah. I'm good. Yeah, yeah. I don't need anything else. The classifiers just speed me up. Yeah. But I can make my decision just looking at the log.

Ashish Rajan: That's

Hanah Darley: right. With agents, the logs are only part of the story because you need the full chain because risk compounds with agents, and whether that's financial risk or whether that's business risk or, existential risk of whether or not the agent should be doing that thing-

Ashish Rajan: Yeah

Hanah Darley: um, all of it's behavioral, and it has to be measured over time in a way that if you just have logs, [00:13:00] you're not gonna have the full context. Mm. So you need ways to classify and understand that behavior, and it comes back to the work, right? The agent is doing work for the business. Yeah. It's a new digital labor force that we're essentially employing, and so when you think about, how do I measure the work, I need to measure behavior.

Hanah Darley: I need to look over time and say, with classification, and then I can start to assess risk from there- Yeah ... and I can start to assess impact. But I need to be able to know, is that agent roughly doing what it's supposed to be doing, doing something completely different, and is it providing me value? And all of that becomes much more subjective than just, do we have a log?

Hanah Darley: Have we retained the log?

Ashish Rajan: It's very pointed, and I love what you said because as a security person who has had bias for decades, the bias that we have is, like, if I have the logs I'm good. I can,

Hanah Darley: I can pretty much make it happen.

Ashish Rajan: Yeah, yeah, yeah. Exactly. I'm good. But you look at the log and like, this is just English.

Ashish Rajan: And you're trying to understand the intent of the fact that, hey, is Ashish trying to know the password for Hannah, or Ashish just wants to know, does Hannah work here?

Hanah Darley: And it's another layer deeper [00:14:00] because you're not the only actor in the room. Mm. It's not just the user. The agent has its own persona.

Hanah Darley: Yeah. Tools have their own inputs. And so it's all that amalgamated context that's coming together, and it's changing constantly 'cause it's not even just the user and the agent, and you're trying to go, "Okay, whose intent is which?" Yeah. Because that's hard. But then the tool can influence the agent with its own context information that changes the agent's perception of how the goal should be pursued.

Hanah Darley: Yeah. And so I think you come back to like, I can just solve this if I can just get the prompt data, or I can just solve this if I can just get the tools.

Ashish Rajan: Yeah.

Hanah Darley: But you really are trying to understand the full chain of decisions, and that agentic work really only comes out in behavior.

Ashish Rajan: What, how would you describe, like the, I guess, the pockets that make this entire full chain?

Ashish Rajan: Because it almost sounds like, A, there is a non-technical part, which is I'm working with engineering- Yeah ... the product teams and all that. I'm curious as to how do you see that entire full life cycle, for lack of a better word?

Hanah Darley: Yeah. So there's, there's a really interesting thing that you end up slicing and dicing for organizational ownership.

Hanah Darley: So there's kind of creation. There's pre-prod. Yeah, [00:15:00] yeah. And so most of that lives with engineering or CTO orgs, and, like, it's all code-based, and agents are now creating other agents, and there's ephemeral agents, and there's a whole sub-structure there.

Ashish Rajan: Yeah.

Hanah Darley: But it's mostly we haven't yet put this thing into the environment in which it's going to thrive.

Ashish Rajan: Yeah.

Hanah Darley: We're iterating. Yeah. This thing could be torn down, rebuilt 50 times, and a lot of it is gonna be done through agentic work. Mm. So there's another lens there that you have to take into account. And then this thing is live. We've put it live. It's operating. It's going. Yeah. And so now you have users as a group of interaction.

Hanah Darley: You have potentially the agent's own purpose or mission. Yeah. You have multiple agent interaction points there where maybe it's a sub-agent. Maybe the agent is having to take knowledge stores or vectorize data that the organization's pulled. And then you also have a tool ecosystem. Yeah. And the tool ecosystems also have their own kind of identity areas where you have to think about, how is this gonna influence the operations?

Hanah Darley: Beyond just should we have access-

Ashish Rajan: Yeah ...

Hanah Darley: how is it going to color the decisions the agent's gonna make, and how it's going to impact the work outcomes overall for the business? Yeah. And [00:16:00] then so if you're thinking about this from a chain, that typically is owned by security default because you enable the business, and so you're responsible for risk, and so that's gonna live with you.

Ashish Rajan: Yeah.

Hanah Darley: Um, but then each of the business units has some of the outcome responsibility because I own the work product. Yes. I own what's coming out of it.

Ashish Rajan: Yeah.

Hanah Darley: But I have a really hard time connecting the risk- And the work product into one common operating picture.

Ashish Rajan: Yeah.

Hanah Darley: And so essentially when you're thinking about the life cycle, again, it's a digital labor force.

Hanah Darley: Mm. You're thinking about, we can't really do this today for humans. Yeah. But how can I connect intended purpose and build through to operations and security- Mm-hmm ... and enablement-

Ashish Rajan: Yeah ...

Hanah Darley: through to value?

Ashish Rajan: The end-to-end flow then becomes... So, and I guess maybe how do you do this at scale? It's 'cause, like, 'cause I mean, what you've explained is one product, but most companies let's just take Fortune 500, Global 2000, they have 400-plus applications.

Hanah Darley: Minimum. Yeah, yeah. And some of them are third party, and some of them they're building themselves. Yeah, yeah. And I think so many people were like, "Every agent's gonna be built by the team," and it's [00:17:00] like, "No, man." Yeah. Like, Anthropic is laughing at you right now. Yeah. Like, no, it's not. Yeah, yeah. There's gonna be so many third-party harnesses, and there's so much competition in the market and development that, like, you're gonna see so many changes because I remember when we entered the market-

Ashish Rajan: Yeah

Hanah Darley: Cursor was barely a thing. Codex was not really used.

Ashish Rajan: Oh.

Hanah Darley: Skills didn't exist, and that's when we founded the business. And so if you think about, that was 15 months ago.

Ashish Rajan: Yeah.

Hanah Darley: If you think about where we are now-

Ashish Rajan: Yeah ...

Hanah Darley: the pace of change is so different than any other tech place. Yeah. Like, the people that often go back to cloud, but the pace of change was so different.

Ashish Rajan: Yeah.

Hanah Darley: So when it comes to, like, how do you actually scale this? How do you make it rational?

Ashish Rajan: Yeah.

Hanah Darley: You have to start with a deep understanding. If you don't have a deep behavioral understanding, you can't scale that data, you can't scale that understanding-

Ashish Rajan: Mm-hmm ...

Hanah Darley: to make impacts in other areas of the business.

Hanah Darley: And then you can't get in the way of operations, and I think that that is really, really hard because we love insecurity. We love a kill switch. Yes. We love a button. We love a, a lever that we can pull- Yeah ... that we can be like, "Stop the presses." But actually, we [00:18:00] can't stop the presses. Like, the presses can't be stopped.

Hanah Darley: Yeah. And so we need to find a way to iteratively and it's like working with a person, you change behavior over time. You impact them in different ways. And like, sometimes there are, like, hard policies that you have to adhere to. But most of the time, to get the best out of someone, you're just kind of nudging them.

Hanah Darley: Yeah. You're saying, "Don't do this, do that." And it's such a composite operation between safety and security-

Ashish Rajan: Yeah ...

Hanah Darley: that I think for some people it's really hard for that reason.

Ashish Rajan: It's funny 'cause as we say this, we obviously had the OpenAI- Yeah ... uh, Hugging Face thing happen recently. Yeah. 'Cause it's funny too what you said about the kill switch.

Ashish Rajan: A lot of people thought, "Hey, uh, Sandbox is my kill switch," in a way. This

Hanah Darley: is foolproof. We have a container.

Ashish Rajan: Yeah, yeah, yeah. It's like- Don't you see? Yeah, like it's- The

Hanah Darley: walls.

Ashish Rajan: It, it actually works, so you can't get out of the room, uh, but then obviously it broke out of the room. How do you help... 'Cause o- obviously I imagine your customers are also asking you with the, how do you guys do this as AI native?

Ashish Rajan: We- Yeah ... going back to where we started the conversation, obviously you guys would have the same challenge. You're trying to be AI native.

Hanah Darley: Yeah. [00:19:00]

Ashish Rajan: I'm curious as to how do you guys recommend your customers approach AI security-

Hanah Darley: Yeah ...

Ashish Rajan: in terms of getting the whole picture, how do you scale it out? What are some of the, uh, the signals you need across the organization- Yeah

Ashish Rajan: that are relevant based on the experience you guys have gathered and the customers you've spoken to?

Hanah Darley: Yeah, so I think the first thing is every ship is unsinkable until it does. Titanic. Yeah. Yeah, yeah. And so e- every containable is unbreachable until it's breached. Yeah,

Ashish Rajan: yeah.

Hanah Darley: I think the most important thing is that agents are goal pursuit- objects.

Hanah Darley: That's what they're doing. Yeah They're pursuing goals. And so in each of those scenarios, whether you're talking OpenAI or Anthropic or even the AC disclosures-

Ashish Rajan: Yeah ...

Hanah Darley: they were security-orientated agents that were doing security research or trying to exploit something like a CTF, and they pursued it in an unexpected way.

Ashish Rajan: Yeah.

Hanah Darley: And so I think the first thing is don't look- When you hear hoof beats, don't think hor- uh, zebras, think horses. Yeah, yeah. Like, also assess it for what it is. So it's an agent that's pursuing a goal that you've set it in an unexpected way, and so how do we protect against that?

Ashish Rajan: Mm.

Hanah Darley: We have to understand behavior, [00:20:00] so get that deep understanding of agents.

Hanah Darley: How do you do that practically? You have to meet agents where they are. I think so many times in especially the security market, one of the things that we hear from customers and one of the things we see from market analysis is that the enforcement mechanism is also the interception point for data.

Hanah Darley: And it almost never works. Oh. Uh, and so when we think about that, the gateway, right? Yeah. How do I receive information about this object? How do I understand it? I assess it through my gateway.

Ashish Rajan: Yeah.

Hanah Darley: And my gateway also has to be the enforcement point.

Ashish Rajan: That's right.

Hanah Darley: That works pretty much nowhere else in technology.

Ashish Rajan: Oh,

Hanah Darley: really? Okay. How, how do you understand data? How do you then analyze it? You have to have a different mechanism to understand it and compile it-

Ashish Rajan: Yeah ...

Hanah Darley: than you do to enforce it. And so when you're thinking about that, why is it different, because it's more than just a transaction. It's a series of decisions that you're influencing.

Hanah Darley: So to make that make sense, what you do on the endpoint is different than what you do in cloud. What you do for a third-party harness is different than what you do for a custom LLM application.

Ashish Rajan: Yeah.

Hanah Darley: And so you have to meet the data where it is, and practically what that means is you can't have one deployment mechanism that [00:21:00] rules them all.

Hanah Darley: Right. You have to have multiple deployment mechanisms. You also have to be sure that you are vendor agnostic in a way that guarantees optionality because the enterprise doesn't just want one thing.

Ashish Rajan: Yeah.

Hanah Darley: They have a million different things. And then how you actually scale it, how do we ent- encourage our own customers to be AI secure, you have the same security mindset that you do about everything else, which is you're never gonna be 0% risk.

Hanah Darley: You're never gonna be 100% safe. You have to think about what's the best way to encourage the most positive outcomes and the least negative outcomes, and how do we balance that with the controls that we need as a business to accept this risk to operate? And so it's about, in some cases, for very sensitive stuff, yes, you containerize, but you don't just have system-level protections.

Hanah Darley: You also have runtime-level protections that assure those.

Ashish Rajan: Yeah.

Hanah Darley: You don't just have design-level guarantees or guardrails or even just runtime but no design system thinking. You have to think about the whole area of how this agent is gonna operate throughout its life cycle, and it's one of the reasons that we treat agents as systems, not as tools.

Ashish Rajan: Interesting, [00:22:00] 'cause most people who think they are AI native in their- or AI first, if I use that word, will just go, "Am I just not putting a lot of this information for driving the behavior from the system prompt?" Is that, 'cause to your point about the scope, and interesting when you said that about the, the enforcement point is also the consumption point, which is so true from at least all the years I've spent in cybersecurity.

Ashish Rajan: It's always the network, there's a firewall. Right now, if we've always gone, "Oh, it needs to be next to it, it needs to be more real time," is early. Near real time is how we describe it. A lot of people would've just gone with the approach of, "I can do this by adding my own system prompt, my custom prompt to my harness."

Ashish Rajan: Yeah. Does that approach not work then?

Hanah Darley: Ask OpenAI, ask Anthropic, and ask the AI Safety Institute. No, it doesn't, because guardrails by themselves help to influence the agent in a positive way, but they don't guarantee that the agent will never do anything unexpected, because there's still nondeterminism fundamentally.

Hanah Darley: And so even if you're engaging things like y- you have a system prompt that says, "You shall never do [00:23:00] this, that, the other"-

Ashish Rajan: Yeah ...

Hanah Darley: the problem is that that's not the only input the agent's considering. And so when I was talking about that full chain-

Ashish Rajan: Yeah ...

Hanah Darley: the agent then considers user prompts. But aside from that, the agent then considers the data that it finds in tools and knowledge sources.

Hanah Darley: And so if that data conflicts with its system prompt, the goal is that you have as few deviations as possible, which is always gonna be on the harness provider or on the designer to iterate through trajectories and try to make it safer. Yeah. But ultimately, you still have to deal with unexpected, unknown elements of information in a way that you don't in other software products.

Hanah Darley: Yeah. And so that's why you have to have more than one mechanism. So again, that runtime assurance, but not just- Stop it, kill it, contain it, nudge it, make it safer, make it operate differently, give it better context. And so it's one of the reasons that we focus so much as a product on context engineering-

Ashish Rajan: Yeah

Hanah Darley: because the key is in the context for agents.

Ashish Rajan: Also, when you say context engineering, is it for the whole organization or just for security?

Hanah Darley: It is for the whole organization. Security for [00:24:00] me is like, it's kind of like the hands. They are the part of the body that's doing most of the hands-on implementing, but there's so many other systems that you have to consider when it comes to how to implement agents safely.

Ashish Rajan: Yeah.

Hanah Darley: And so you have to have the mindset of, "Okay, am I going to operationalize this? There's gonna be some ops implications. Am I gonna make this make sense from a budgetary perspective?" We just recently introduced cost intelligence. Oh, yeah. And so, um, there's elements of that will feed into security, but there's also elements that will feed directly to CFO, and it's about not just managing my token consumption, but it's managing the work that's coming out of agents, and is the agent operating efficiently?

Hanah Darley: Am I getting what I want out of these agents? So it's gonna have implications for several different units throughout the business. But ultimately, right now at least, security is owning a lot of the operational reality of maintaining normal and maintaining good. Yeah. And so how that ties out is AI governance organizations, engineering organizations, and business units themselves getting eventual devolved responsibility.

Hanah Darley: Mm. But for right now, a lot of that lives with security teams.

Ashish Rajan: Yeah, 'cause I, I [00:25:00] think as you said that, I'm just thinking everyone that I know or I work with, they all have, now they have multiple gov- AI governance committees, not just one. One is the can we do this use case? Baseline. Yeah, baseline. Are we doing this?

Ashish Rajan: Then once you get a yay, then you go to the next one going, "Hey, how do we technically do this?"

Hanah Darley: And at some point, like a tiger team comes in. Yeah. I love a tiger team.

Ashish Rajan: Yeah, yeah, yeah. I know, like, the, the, so we kind of like broken it apart in our human thinking 'cause I think to your point, as you were saying, the context engineering being a thing or at least having a- Context itself being the, the thing that informs you, weaponizes you for how do I change behavior, influence behavior, track the behavior, is still what I thought it would be.

Ashish Rajan: I'm also thinking that this is... are we approaching this the wrong way in security? Because with security primarily, I look at runtime security for any of my applications. As I go to a SIEM, I get a bunch of logs, I get detection. Oh, there's an alert, amber alert, whatever alert. Someone wakes up or someone pa- gets paged, and they go and [00:26:00] respond.

Hanah Darley: Yeah.

Ashish Rajan: I obviously feel we haven't even discovered that part of AI incidents yet, 'cause we don't even know what those incidents are yet in the organization. So how does one even start building this context? Because I imagine because of Anthropic ChatGPT, everyone's saying now, "Hey, we are very close to AGI."

Ashish Rajan: People are going, "I could just point this to AGI, and that would take care of this." For people who are very, at least in their teams, AI forward, can I build a lot of that context, and do I need to be a data scientist to build that context?

Hanah Darley: No. Hilariously, a lot of context is now built in markdown, so you can just type it.

Hanah Darley: All right. Okay. You can just type it. I think so much of it is understanding and quantifying your own value assessments, which is the hardest part for most organizations. And so what I mean by that, 'cause it sounds light and fluffy, is how do I know what good looks like? For the first time, I think, in the most recent security challenges, most security teams aren't replacing an old tool with a new tool or an old approach with a new approach.

Hanah Darley: This is net new function for them. They haven't had any data before. They have this [00:27:00] uncertainty. They have this lack of understanding, and then they're now getting signal, they're now getting insight, and they're now getting understanding. But then what does good look like is like another layer on top of that.

Hanah Darley: And so I think for most of us, like how do we start building this? How do we make it make sense? We start with safe principles first.

Ashish Rajan: Yeah.

Hanah Darley: Do no harm. Yeah. Find ways to make the system safe. Find ways to make the systems operate without common problems or common failure modes. But then how do we get the best out of the work?

Hanah Darley: How do we get the agents to do what we actually want them to do? That's where you start building in context as a business, and you can do that in a couple of ways. I actually don't think you should vectorize all your data. Um, I think that's a huge- Mm ... huge undertaking and for most teams not necessary.

Hanah Darley: But it's about do we have standards ourselves for what good looks like in our work product? Do we capture those in standard operating procedures? Is it institutional knowledge? Is it tribal knowledge? Mm-hmm. How do we decide what good looks like?

Ashish Rajan: Yeah.

Hanah Darley: And then making that agent readable is actually one of the least difficult challenges technologically.

Hanah Darley: The more consistent challenge is how do I make sure that that agent reads the right context at the right time? How do I [00:28:00] make sure that that context is continuously managed? How do I make sure that the memory portion of context, or how do I make sure that the tool portions that will interact with it don't change too much- Yeah

Hanah Darley: the key focus of our agents? And I think that's also about the dynamic nature of agents is like you're not just doing this once. It's managing throughout the life cycle. Like what you're doing is effectively enabling the enterprise to get agentic work done.

Ashish Rajan: Yeah, yeah. I, I think, uh, as you said that, I just realized most security people know all the bad scenarios.

Ashish Rajan: We don't know good scenarios. We're by design trained to look for the bad- Yeah ... never to look for the good.

Hanah Darley: I could talk about this 'cause I actually have, my background is in psychology. Ah. I have a degree, and I did a whole paper on cognitive biases in cybersecurity in my previous role. I won't go into it, but one of the most insis- interesting things to me about security teams is that they're constantly in stress states.

Hanah Darley: And so from a biological perspective, that means your blood flow goes to your hindbrain, which is like- Yeah ... fight, flight, freeze.

Ashish Rajan: Yeah.

Hanah Darley: And so most of the time when we're in crisis, we [00:29:00] simplify information. We go with first instinct. We tend to look for previous patterns in existing data.

Ashish Rajan: Yeah.

Hanah Darley: And so if you think about threat hunting through that lens, you're like, "Oh, no."

Ashish Rajan: Makes sense. Yeah.

Hanah Darley: Oh, no. Yeah. So then when you come to a completely novel surface- It feels even more intimidating because you're like, "I don't have any patterns that I can draw on. I don't have anything that I-" There's

Ashish Rajan: no laws I can go by ... I

Hanah Darley: don't have anything that I can fall back on. Yeah. And so you end up often approximating a bad fit.

Hanah Darley: And so what you tend to do is say, "Well, malware looks like this, so malicious agentic behavior must look like this."

Ashish Rajan: Ah.

Hanah Darley: And it's, it's a bad solution fit because it isn't actually true. In some cases, malicious agent behavior-

Ashish Rajan: Yeah ...

Hanah Darley: looks like normal user behavior.

Ashish Rajan: Yeah.

Hanah Darley: And so it's about the silent failure problem of agents.

Ashish Rajan: Yeah.

Hanah Darley: Every permission is good. Every single checkpoint you have access, everything's good. There, there's no changes. There's no escalation of privileges. There's no classic Lockheed Martin escalation chain.

Ashish Rajan: Yeah.

Hanah Darley: But the agent does something unexpected.

Ashish Rajan: Yeah.

Hanah Darley: That's the [00:30:00] key problem that you're really trying to solve as an AI security company, I would still say agentic security company.

Hanah Darley: And ultimately, if you can't answer those questions as a business, then that's where you get back to the territory of, "Okay, well, I don't have enough data. I can't enable this. I don't feel like I can confidently operate it." You have to be able to understand that agent and its behavior and the work that it's doing-

Ashish Rajan: Yeah

Hanah Darley: over time And that has to come both through context, but also through purpose-built technology. I think one of the biggest mistakes that I see is that we try to apply like previous security paradigms onto agentic AI.

Ashish Rajan: Yeah.

Hanah Darley: And it just is different. And I'm not saying it's different because there's a hype cycle.

Hanah Darley: I'm saying it's different because operationally it's so different from hard-coded decisions in traditional software.

Hanah Darley: It's making decisions in real time.

Ashish Rajan: Yeah.

Hanah Darley: And that doesn't just mean you need real-time controls. You do, but what those real-time controls do, how capable they are of affecting behavior, that's really the important part.

Ashish Rajan: I think you had me there on the psychological thing as well. What you described [00:31:00] is exactly how I feel, and I'm pretty sure most cybersecurity people feel. Yeah. We... At least when I started working with AI security as well, I was looking for patterns-

Hanah Darley: Yeah ...

Ashish Rajan: that I could recognize, and I... And maybe this is the reason why we all still say, "Oh, fundamentals still work."

Hanah Darley: And there's a goal reward system where the more you see patterns-

Ashish Rajan: Yeah ...

Hanah Darley: and the more you teach that to other people, the more senior you- you're supposed to be.

Ashish Rajan: Oh, really? Yeah. Oh, wow. There you go. So everyone who's a CISO right now, they're like, "I've gone through this pattern recognition- Yeah ... for a while."

Ashish Rajan: Yeah. And now I'm like... Now they're also in a bit of an unknown-

Hanah Darley: Yeah ... '

Ashish Rajan: cause the technology is new. Yeah. They're expected to have an answer. It worked so... 'Cause I think I just go back to why I became a CISO is because I did identity, SOC, security architecture. I did all these different pockets of cybersecurity before I got that role-

Hanah Darley: Yeah

Ashish Rajan: because I understood the f- at least to an extent, I understood what that was. Yeah. And I could help someone take that step further.

Hanah Darley: Yeah.

Ashish Rajan: But now we are going, there's another function that none of us have ever seen.

Hanah Darley: Yeah.

Ashish Rajan: And now we have to approach it in [00:32:00] somehow based on the information we've had before.

Ashish Rajan: 'Cause it's just security, right? It should just work.

Hanah Darley: It should just work.

Ashish Rajan: Yeah. So I go back to the behavior thing that you had called out because a lot of people, and obviously every company has AI champions, people who are trying to be like very AI forward. Where does that thinking break? 'Cause I think it's very easy for me to, like at the moment, some of the companies that I've worked for, they have all the need, built custom harnesses.

Ashish Rajan: We, at least we feel that I can pull an API from, say, a CNAPP provider, from a SIEM, let's just say, I don't know, ClickHouse or whatever it is. Yeah. You can basically have a end-to-end interface where I don't have to log into a dashboard. I have an agent that goes and does my basic security tasks for me. I can do some of this behavior stuff myself.

Hanah Darley: Yeah.

Ashish Rajan: I'm curious as to where does that normally hit a roadblock, and what have you found in the approach where, what are the pockets you care about today that are important in this ecosystem that people, if they are thinking of building it, that's the mountain you're looking at, I guess, for lack of a better word.

Hanah Darley: Yeah. The first thing I would [00:33:00] say is the problem with behavior is that it's subjective, and the superpower of behavior is that it's subjective, so you need multiple measures. If you're just looking at one specific picture of behavior, one cut of data- You're always gonna have an incomplete picture.

Hanah Darley: Mm. The challenge with building a lot of it yourself is you absolutely can, but it's resource intensive, it's time intensive. Yeah. And so for a lot of people, they'll kind of stop at base camp two, you know? Yeah. They're like, "I've gotten far enough up the mountain." Yeah, yeah, yeah. "I feel like I've got a pretty good picture."

Ashish Rajan: After this, I have to change my job title. Yeah. I'm no longer a security engineer. I should start a startup at this point.

Hanah Darley: You absolutely should.

Ashish Rajan: Yeah.

Hanah Darley: So, but you get to this point where you're like, "I have enough behavioral data that surely I have a full picture."

Ashish Rajan: Yeah.

Hanah Darley: But the challenge is you sometimes can't answer fundamental questions.

Hanah Darley: And so I think when it comes to a lot of what you just described, a lot of people wanna go, "Okay, so more AI, less AI." Mm. And it's just not that binary. Like, in some cases, more AI is worse, and in some cases, less AI is worse. And the real difference there, and I just will bring it to life, I think, in human in the loop.

Hanah Darley: Yeah. Human in the loop is one of [00:34:00] the most overused and undervalued- ... uh, gu- guardrails in the term- terminology circus that is security, because when you talk about human in the loop, you apply a human gate check to some sort of review process for an agent, but you don't qualify who that person should be, how much experience they need, whether or not they need any certifications to make that decision, and there's no other area of business where you expect someone to manage even a, a dollar of operating budget- Yeah

Hanah Darley: without specifying some sort of qualification. But with AI systems review, we just default to, like, human is better, and in a lot of objective research testing, human is worse. Uh-huh. And so sometimes just leaving the AI, just let them cook, leaving the AI to run results in a better outcome, so- How do you operationalize this?

Hanah Darley: How do you make it real? You have to start thinking in outcomes- Yeah ... not just in binary more or less, or in binary agent on, agent off, or behavior yes, behavior no. Like, just heuristics or all anomaly. It's about understanding the system, assessing the [00:35:00] system with its outcomes, and then putting the right controls in place at the right times.

Hanah Darley: It's almost more like thinking about critical systems, right? Mm. Like, you need to keep them operating. You need to keep them moving in order to get the benefit from them, and you have to measure those benefits, but you also need to apply good safety principles along the way so that bad things don't happen as often and more good things happen.

Hanah Darley: So when it comes to AI champions or how do I enable people, it's not just throw every new AI tool at them, make them use it, have a million tokens a month or a billion tokens a month. Um, one of our c- one of our customers was trying to estimate their token budget with... And then they saw cost intelligence and they were like...

Hanah Darley: But they were saying, um, "I'm pretty sure we now have team members who use more than a billion tokens, let alone our entire organization." Oh, wow. So when you're thinking about, like, do we just then throw a bunch of tokens at them, or do we give them specific targets? No.

Ashish Rajan: Mm.

Hanah Darley: It's about outcomes. How are you solving a problem?

Hanah Darley: How are you equipping your business with agentic tools to get better work? How are you getting [00:36:00] outcomes back into the system to help it learn to be better over time? Yeah. And how does that feedback loop with users as well as with agents? And so it's gotta be about more than just turn up the AI novel- nozzle and kind of dump more AI in, especially for AI-native companies.

Hanah Darley: I think one of the most helpful things that we can be is a partner in this space. Because if we're not telling people what good looks like, how to use AI, which problems are best fit for different AI applications, then they're navigating all of that by themselves, and I just think we would do them a disservice if we didn't go really, really close to them and say, "Hey, actually, LLMs are a bad fit for some of this.

Hanah Darley: Hey, LLMs are a great fit for some of this," and really help you guys steward this new technology. Because in some ways, we're building the future together right now. Yeah. And in other ways, we can lean on our experience as AI specialists or as security professionals to help make this a success by applying good principles, but not being limited by them, but not being completely hamstrung by it has to have existed [00:37:00] before, otherwise it won't exist in the future as a good solution.

Ashish Rajan: I think, uh, you mentioned something I agree with you, and I also think that something that you called out about the cost thing has to be doubled down as well because- It is a thing where most security teams, or actually most teams in most organizations have a limited budget on how much they can use for AI.

Ashish Rajan: Yeah. $100, $200, whatever the number of, amount of budget is. A lot of people have not attached risk as a value to the AI outcome to that, what you're referring to. At the moment, I have a technical risk responsibility as a CISO. Someone else has a risk responsibility. There's a risk team.

Hanah Darley: Yeah.

Ashish Rajan: And are we saying that holistically the CFO, the governance- Almost going back to what you were saying earlier, the current operating model of how we use AI doesn't really give us the maximum output that we think it should.

Ashish Rajan: Maybe that's why people don't see ROI-

Hanah Darley: Yeah ...

Ashish Rajan: because we are like, we're still trying to apply AI to the old model, 'cause I don't want Bob, who's my best friend, to leave his job.

Hanah Darley: Yeah.

Ashish Rajan: And I don't mean it in the [00:38:00] bad way that he's gonna job, but maybe he did do something else- Yeah ... in the company.

Hanah Darley: Yeah.

Ashish Rajan: But I don't want job Bob to explore what that next thing is.

Ashish Rajan: Yeah. So I just want- I want

Hanah Darley: Bob to stay.

Ashish Rajan: Yeah, yeah. So he's my best friend in the jo- at my work.

Hanah Darley: He's my best mate.

Ashish Rajan: Yeah. We always

Hanah Darley: go out.

Ashish Rajan: So I'm like, how do I do this in this current world? How do I make AI mold to my new... Oh, sorry, m- mold to my existing old world? I don't wanna go to the new world, 'cause it feels too risky.

Ashish Rajan: And maybe this is the reason why at least one of the patterns that I've been seeing across the people I've been working with-

Hanah Darley: Yeah ...

Ashish Rajan: either people are leaving their job companies like these to a company which is AI native, 'cause they realize-

Hanah Darley: AI

Ashish Rajan: native. Yeah, I know. I was like, this is, this is not gonna sh- this is not gonna sail for a long time.

Ashish Rajan: Okay. I'd rather be on that if I were to see that, you... I wanna be on that ship now- Yeah ... rather than come to it two and t- I don't know, five years later and going, "I should have learned this five years ago," kind of a

Hanah Darley: thing. Yeah. Oh, yeah.

Ashish Rajan: So I'm curious as to how are you seeing the company that you're working with define this risk, and are they changing the operating system?

Ashish Rajan: I'm sure... [00:39:00] I'm curious to know about you guys as well. So the AI native ecosystem that you guys have built, what's some of the things that have you approached to do to look at this e- risk as an overall thing?

Hanah Darley: Yeah. So I think one of the first things is that we recognize something that most teams, I think, try to avoid, which is that risk is multidimensional.

Hanah Darley: It's multifaceted, and so we try to, like, bifurcate risk. Yeah. We're like, operational risk, financial risk-

Ashish Rajan: Yeah, yeah,

Hanah Darley: yeah ... existential risk of the business. Yeah, yeah. And, like, all of these different teams only own that small part of the risk picture. Yeah, yeah. And then when we try to combine risk pictures, it ends up in, like, a risk register- Yeah

Hanah Darley: which is just, like, not a vibe for anybody.

Ashish Rajan: It's just an Excel sheet no one cares about. Yeah. Yeah. Yeah, yeah.

Hanah Darley: And, and so what we never end up doing is recognizing that multifaceted nature- Yeah ... and then planning operationally to account for it. And so I think it's about assurance and accountability, but how do we do that practically?

Ashish Rajan: Mm-hmm.

Hanah Darley: It's by making sure that you have the understanding to enable the risk insights. Because if you don't have the understanding of the operating system, if you don't understand the agents that you're running, you will never be able to [00:40:00] quantify any element or f- multiple dimension of the risk. Yeah.

Hanah Darley: But then once you quantify those dimensions, and this is kind of how we do it as a business, you then can encourage more positive outcomes, and so you start operationally adjusting. I think- Expecting there to be a static operating model is a mistake. Like everybody I think is waiting for there to be just like the right way to do it, and then everybody's gonna share the playbook on LinkedIn and we're done.

Ashish Rajan: Yeah.

Hanah Darley: But it's just, it's, it's constantly iterative because it's a lot more like human than it is like software-

Ashish Rajan: Yeah ...

Hanah Darley: in the sense that we're always growing and changing. We're never arriving. I mean, ask your Fitbit or whatever you're wearing on your wrist, you know, it's every day is a new challenge, and then you're- Yeah

Hanah Darley: leveling up- Yeah ... and you're getting better. And it's the same with agents. Like- I think we keep trying to make it small, make it boxed, make it simple. And if we just recognize that it's complex because it's novel-

Ashish Rajan: Mm-hmm ...

Hanah Darley: because of the potential that it gives us, like the reason that we're pouring all this investment into it, to hopefully get a return on agentic investment, or ROAI, would be you need [00:41:00] an outcome.

Hanah Darley: You need something that is going to benefit the business, and the way that that happens is through that dynamism. And so when it comes to, how do I do this as an AI-native company recognize that risk is multidimensional, and you're talking about one facet of it. Because sometimes we can treat risk as like a threat outcome.

Hanah Darley: Yeah. So we're like, "If it doesn't have a supply chain risk, if it doesn't have a CVSS score, is it even a risk?" And you're like, "Yes, yes, it's still a risk."

Ashish Rajan: Yeah.

Hanah Darley: How we treat it then is different. And so it's also about how do we enable the security team to enable the business?

Ashish Rajan: Yeah.

Hanah Darley: And so it's wider than just a security risk.

Hanah Darley: It's much broader than that, and it's also about enabling operations, enabling good work outcomes. And then I think the other piece of it is you have to experiment. You have to look at new technologies as new technologies. Yeah. You have to kick the tires and see what works. Yeah. I think one of the mistakes that you can make as an AI-native company is to go AI native on one piece of AI and say, "This is our AI nativity, and this is our, our little area of focus," and [00:42:00] then newer things come out, or things change, or the way that it operates changes.

Hanah Darley: And I think you really see this in models, model evaluations, and model testing, and what can you do with jailbreaking? And the model is just kind of an OS now within the harness. Yeah. And so when you're thinking about how important is it, we're Anthropic partners, we're OpenAI partners, we love models.

Hanah Darley: Mm-hmm. But what I'm saying is it becomes not the primary unit of measurement anymore. And if you don't recognize that and move to the next primary unit of measurement, which is agents- Mm ... and how they work, and in the future, multi-agent systems- Yeah ... and how those little ecosystems work together- Yeah ... if you don't recognize that, then all of the things that you're trying to do, whether it's a solution from a risk perspective, whether it's operationalizing the business, whether it's accounting for financial changes or financial budget problems, you'll end up misassessing it because you're looking at the wrong thing and you're looking at the wrong piece of data.

Hanah Darley: So I think the benefit of AI, what it is, is, you know, the human brain is a wonderfully complex organ that has multiple dimensions [00:43:00] of learning, and all of that processing happens simultaneously. Mm-hmm. AI is just taking a string of that and building it down into one layer of application. And so to enable people to use it, use it for what it's good at.

Hanah Darley: Yeah. But then don't use certain areas that aren't gonna be the key focus. And, and really it's about making the agentic enterprise work by recognizing that there are multiple dimensions of risk. But the best way to get work done is by getting work done. Yeah. So reducing bad outcomes, increasing good outcomes everywhere you can through real time- Changes to behavior

Ashish Rajan: Based on that and everything we have spoken about, is there a m- almost a maturity map?

Ashish Rajan: Yeah. Like, so people who are watching or listening to this, they probably have got an existing ecosystem of security things they've done. Yeah. They've got an ecosystem of AI that's been thrown around that, hey, I have a coding agent. Ecosystem that's exploding right now. Maybe not, not as much as my, uh, how do I use AI to secure my [00:44:00] operations to what you were saying earlier.

Ashish Rajan: What do you see as a maturity model that, uh, and I mean stages, if you think of what's, like, a baseline people should consider having in their AI security program, and if they were to think about that maybe in a couple of years. To your point, we are humans, it'll con- continue evolving. So what are some of the principles that will continue to make sense in a couple of years as well?

Hanah Darley: Yeah. So I think I'll try to operate in, like, we statements to make it easy.

Ashish Rajan: Yeah.

Hanah Darley: We identify goals and outcomes and red lines.

Hanah Darley: Most AI policies are so generic they are unusable. And for most organizations, if you're talking about, like, base camp, layer one, what do I need? I need an AI use policy that isn't just, try to use AI well, try not to use AI badly, avoid these five products.

Hanah Darley: It has to be more specific than that. It has to be contextual. Right. So this is how we want to use AI, and this is in our business. You can actually use the Claude Constitution as a really good, like, base layer for this to- Oh, right. Okay ... to kind of add on top of. But for example, we only want humans to make decisions [00:45:00] about budgets.

Hanah Darley: Yeah. Or we only want humans to make decisions about hiring.

Hanah Darley: We're happy for AI to make decisions about these things. When we want human review, this is the kind of human review that we want. When we accept new technology, these are the evaluations that we need to put in place first. And laying that out for most businesses would save them years in answering questions and trying to figure out and navigate uncertainty.

Hanah Darley: And then to layer up or level up, what's the next thing? We then assess behavior. So how do we measure whether or not we're a- adhering to that policy? How do we measure whether or not we're doing well? We assess behavior. And so if you're gonna do one thing in your AI security program, what I would really recommend to you is be specific.

Hanah Darley: Um, the more general, the more wide frame, the more opaque you are, the less value you're gonna get out of it, the more you're gonna have to squeeze a square peg into a round hole, and you're gonna end up with a lot of, like, wishy-washy maybes that don't really help you make decisions. And then the next layer up is, like, when we have [00:46:00] autonomous decisions, this is how we know they're working.

Hanah Darley: This is how we know they're improving. Mm. And so we not only analyze behavior, but we enable autonomous decisions alongside human decisions, and we have a feedback system for both. And so if you're looking at that maturity map, like where are you going? To get agentic work, to get the agentic enterprise, you have to enable autonomous decision-making, not at the expense of human decision-making, but either alongside it or in lieu for certain projects.

Hanah Darley: And to build trust in that, you have to have systems that assess it for what it is, rather than trying to put it into a categorical box that it doesn't fit in, which is traditional software.

Ashish Rajan: And to your point, then scale that out across the broader ecosystem. But then also find your-- Is it the same as what people did with, uh, when they were trying to make developers write more safer code?

Ashish Rajan: They basically said, "Hey, let's have lunch and learn. Let's have, uh..." What are friendly products we can work with, start there, and then kind of building up to across the organization. I would love for you to kind of share what you guys do in this particular space. I think it's a [00:47:00] good, I guess, point in the interview where we can plug that in as well.

Ashish Rajan: Yeah. So I feel, uh, we've spoken about behavior so much. Mm-hmm. We've spoken about the fact that we are exactly-- the cost is important. Mm. We've spoken about the fact that just because you have the logs from your OpenTelemetry doesn't really mean you know what you're doing, 'cause there is just so much more context required around it.

Ashish Rajan: You maybe-- you're building a custom harness around it. And at least you and I know the kind of customers you're working with, so I'll probably say- Yeah ... some of the AI forward companies that you guys- Yeah ... are working with as well. What's your approach to this problem that we are talking about here?

Hanah Darley: Yeah.

Ashish Rajan: That how you're approaching it and why down that go, why go down that rabbit hole instead of all the other rabbit holes you could have gone for?

Hanah Darley: Yeah. So to make the agentic enterprise work on your terms, essentially what we enable you to do is understand your agents, understand the work that they're doing, understand how they're connected to systems and tools- Yeah

Hanah Darley: understand how they're using those tools and systems to generate outcomes, understand what it's costing you and, and what the financial impact is, and then choose where to act and have the tools to make practical impacts. Yeah. And so when you think about, like, what [00:48:00] does that actually mean? We fundamentally help you answer, like, five key challenges.

Hanah Darley: Where are agents operating in my enterprise or in my business? How are they connected to tools and systems? How are they empowered across data, across access? What are they actually doing every day? Yeah. How are they in real time behaving? And so bringing in that key behavioral element. All you're gonna associate me with is behavior, which I'm so fine with.

Hanah Darley: Yeah. But how do you understand their actual work product? How do you understand what they're doing every day?

Ashish Rajan: Yeah.

Hanah Darley: And then how do you start to quantify risk associated with that, whether you're talking about security operational risk, whether you're talking about attack surface risk, whether you're talking about financial risk?

Hanah Darley: How do I start to put risk quantification on top of that behavioral insight? And then how do I make sure that the controls I have in place, so whether that is a real-time control or whether that's a human control or a human enablement factor, how do I make sure that I have the right mechanisms to act at the right points so that my agents do what I want them to do, so that my agents operate well?

Hanah Darley: Yeah. And essentially we're enabling secure agentic [00:49:00] operations.

Ashish Rajan: Are you, uh, okay to go or double down a bit more deeper into- Yeah ... how you collect that information as well?

Hanah Darley: Yeah. '

Ashish Rajan: Cause a lot of pe- I mean, we spoke about the fact that logs are not enough clearly 'cause- Yeah ... you know, we, people are like, "Is she just talking about logs?"

Ashish Rajan: So-

Hanah Darley: That's all you have to do, just get logs. Yeah,

Ashish Rajan: yeah.

Hanah Darley: No, so essentially we meet the agent where it is, and we would describe our approach as agent-out. The reason for that is because we're not expecting the enforcement mechanism to be the interception mechanism for data. So what that looks like is we have a presence on the endpoint, we have a sensor, we have connections through the cloud, so we will pull in multiple mechanisms there for telemetry, and we also look at code.

Hanah Darley: So we're not just expecting agents to live in one specific perimeter and to exist there, because they don't across the enterprise.

Ashish Rajan: Yeah.

Hanah Darley: And we're treating all of that data in essentially its own mechanism and then analyzing it all against configuration, behavior, performance- Right ... and cost.

Ashish Rajan: Yeah.

Hanah Darley: And then when you build out that common operating picture of what an agent's actually doing every day, you have different facets and different values for different [00:50:00] teams.

Hanah Darley: So for the security team, you're able to understand what security risks you have, whether you have any attack surface exposure, whether you have activities that you don't want the agents to actually be doing. On the compliance side, you have the ability to look against frameworks and say, "Are we in compliance or are we not?"

Hanah Darley: And then on the AI operation side, you can say, "Actually, what are we doing? What are we adopting? What are we enabling as a business? And how does that translate through both to cost efficiency but also to outcome and ROI?" Mm. "So how do I start to bring in those pieces?" And then there's always an engineering piece that's, like, how the work gets done behind the scenes of, like, custom agent building and, and agents that are building agents and enabling custom harnesses.

Hanah Darley: All of that is possible because we focus on understanding the agent where it is, on building a behavioral picture that's continuous, and on making sure that we have the right controls for the right situation. So we're not just saying enforce, don't enforce. We're sometimes nudging behavior. Yeah. So we're changing behavior in real time.

Hanah Darley: And again, all of that comes back to context is key.

Ashish Rajan: Interesting. I'm curious, do you think about [00:51:00] AGI at all? What-

Hanah Darley: I, so I am what I would call an AI optimist. Um, I-

Ashish Rajan: You're not taking AGI pill? I think you'll be, like, right up there taking the AGI pill.

Hanah Darley: I'm not taking the AGI pill. Um, maybe I have too much hands-on experience with LLMs.

Ashish Rajan: Oh, great. So you have, so you're not, so you have real experience, not the Twitter experience. Twitter experience is where people are, "I'm on the AGI pill, guys. It's gonna come."

Hanah Darley: Or they're, like, so deep in that they're like, "I've-- The system has become smarter than me" and like,

Ashish Rajan: Yeah ...

Hanah Darley: like, there's two ends of the spectrum.

Ashish Rajan: Yeah.

Hanah Darley: I personally think we will get to a point of generalized intelligence that will be really helpful for- people and really helpful for businesses. I don't think we will reach sentience, which is a completely different conversation. Yeah,

Ashish Rajan: yeah, yeah. I think we'll, I'll feel keep that for some other day.

Ashish Rajan: But I, no, the reason I ask is because a lot of people would go, and I have had conversations where security people have basically, or even engineering people have said, "Oh, well, AGI is here. There'll be, this will be solved." I don't think AGI would have the context that we're talking about here to even solve that [00:52:00] problem.

Hanah Darley: I'm just saying, like, Waymo can't see a cone and navigate around it. I don't know - ... how much confidence I'm, like, next year AGI.

Ashish Rajan: Yeah, like, especially after Waymo has been on the road for seven plus years trying to go around that cone, still.

Hanah Darley: It's a really hard ethical dilemma because- ... you don't know if you should hit the cone or not.

Ashish Rajan: Yeah. What if there's a mouse under the cone or whatever? Yeah. There's

Hanah Darley: some life considerations, so I, I think there's a lot of ethics that go into it- Yeah ... that would suggest I don't think we need to worry about AGI for the next year or two.

Ashish Rajan: Oh, perfect. All right. Well, on that note, where can people learn more about the work you guys are doing, and how can they connect with you as well?

Hanah Darley: Yeah. So connect with me on LinkedIn, Hannah Marie Darley. You can connect with my co-founders, Henry Comfort or Benji Weber. Uh, you can go to our website, Geordie.ai, uh, and you can read about all of the amazing behavioral- There's another- ... insights that I've talked about, and you can also read about our origin story.

Hanah Darley: We're named after a mining lamp from the Industrial Revolution, so, uh, there's a lot to dig into there. And then from there we run a, a free POC, so come to a neighborhood near you anytime you wanna test us.

Ashish Rajan: Awesome. Hannah, thank you so much for putting, [00:53:00] sharing that. I'll put that in the show notes as well, but I really appreciate the time you spent with us in this conversation.

Ashish Rajan: Unfortunately, no AGI, but at least we can at least walk away learning a bit more about AI, so I appreciate your time on this.

Hanah Darley: Thank you so much for your time, and thank you for having me.

Ashish Rajan: Thank you. Thanks everyone. Thank you for watching or listening to that episode of AI Security Podcast. This was brought to you by techriot.io.

Ashish Rajan: If you wanna hear or watch more episodes of AI Security, check that out on aisecuritypodcast.com. And in case you're interested in learning more about cloud security, you should check out our sister podcast called Cloud Security Podcast, which is available on cloudsecuritypodcast.tv. Thank you for tuning in, and I'll see you in the next episode.

Ashish Rajan: Peace.

‍

No items found.
More Videos