Traditional two-week manual pentests are no longer enough to protect against adversaries operating at machine speed. In fact, relying on them to cover your entire attack surface borders on negligence. In this episode, Ashish sits down with Travis Lanham (CTO) and Evan Peña (Chief Offensive Security Officer) from Armadin to discuss the reality of AI-driven "Hyperattacks". They explain how they use a swarm of autonomous AI agents, recently unleashing 26,000 agents and 40 billion tokens against a single environment to discover critical vulnerabilities that human testers simply don't have the time or scale to find. We speak about the architecture of an autonomous attack, detailing how "Overwatch" agents coordinate specialized "Hunter" agents to dynamically execute complex kill chains without hallucinating. and why models like Kimi K3 and GLM are far more dangerous in the hands of threat actors than restricted frontier models like Mythos.
Questions asked:
00:00 Introduction & The Reality of Machine Speed Attacks
02:00 Evan & Travis's Backgrounds (Mandiant & Google)
04:00 Why Traditional Pen Testing is Obsolete
06:00 Inside a Hyperattack: 26,000 Agents & 40 Billion Tokens
08:00 Scripts vs. AI: Why Autonomous Agents Beat SOAR & Nuclei
10:00 Building Safety Infrastructure for AI Red Teams
13:00 Agent Architecture: Overwatch, Scout, and Hunter Agents
15:00 Verifying Kill Chains to Prevent AI Hallucinations
18:00 Discovering Critical RCEs Human Testers Missed
24:00 Why Open Models (Kimi K3, GLM) Are More Dangerous Than Mythos
28:00 The Human in the Loop: Blue AI vs. Red AI
34:00 The Future of Detection & Response (Continuous Attacking)
39:00 The "You Laugh, You Lose" Cybersecurity Joke Challenge
Travis Lanham: [00:00:00] 13,000 concurrent attacks, 26,000 agents, 40 billion tokens Now I'm too scared to find out stuff that I don't want to know about. The, uh, new attacks that are shutting down water infrastructure, they're AI powered.
Travis Lanham: Like, the threat actor is AI.
Evan Peña: Very capable models and very violent in nature in terms of, like, they don't refuse anything that you wanna give it.
Evan Peña: I don't know why someone would do a traditional pen test at this point.
Travis Lanham: Today, we're at machine speed, machine scale, machine sophistication.
Evan Peña: We'll find more in that one 48 hours than people will find in an entire external pen test done with a two-week, you know, two-person kind of team.
Travis Lanham: What does it look like to have that relentless swarm coming after my organization?
Ashish Rajan: If your network is attacked by an AI agent, let's just say a swarm of AI agents. By the way, I hope it doesn't happen to you, but if it was to happen, would you want to know what would happen in your organization if there was a swarm of AI agents that attack you at the same time, both from a red team perspective and for vulnerabilities yet you may [00:01:00] not have seen?
Ashish Rajan: Nine out of ten times, you probably wanna be proactive about it. A lot of our conversation these days about pen testing usually is they... "I do it once a month, every quarter." Whatever the sequence of events is, what triggers a pen test in your organization. I had a great conversation with Travis and Evan from Armadin, who have been working on the autonomous pen testing for some time, and we spoke about some things about what's it like to have a hyperscale attack simulate in your environment, why is there a need for automated pen testing more than ever today, and we also spoke about why coverage of testing needs to be more than just a short window that you provide to a human tester, where the advantages are, and how your team, both red team and blue team, can prepare for it.
Ashish Rajan: All that and a lot more in this episode of the podcast. If you are enjoying these episodes and are here for the second or third time, I would really appreciate if you take a quick second to drop the follow or subscribe button, whichever podcast platform you listen or watch this on. We are on Apple, Spotify, YouTube, and LinkedIn.
Ashish Rajan: I hope you enjoy this conversation with Travis and Evan. I'll talk to you soon. Hello, welcome [00:02:00] to another episode of the podcast. Hey, guys. Com- Thanks for coming on the show.
Travis Lanham: Thanks for having us.
Ashish Rajan: I am looking forward to this conversation. Maybe, uh, we'll, we'll start with yourselves. If you can get a bit introduction about yourself, your professional background, where you guys are now.
Evan Peña: So I'm Evan Pena. I am a founder, and I'm our chief offensive security officer at Armadin. Uh, and so what I, what I run is our expertise team, and so that includes threat intelligence, remediation, red team, and operational technology or industrial control system specialists.
Evan Peña: And, uh, prior to starting Armadin I was the global red team lead at Mandiant, where I led that team. I was there for 13 years, and when I started there, we were just an incident response company for the most part. And so I kind of helped build out that entire practice of a, a large portfolio of practice security services on the professional services side.
Evan Peña: So a long era of, uh, expertise behind me and a lot of really cool stories that came with it. I'm surprised you still have black hair. What about you, Travis?
Travis Lanham: I'm Travis, our CTO. Before this, I was the tech lead for Google's enterprise security portfolio around [00:03:00] security operations. And then at Armadin, I focus on, you know, working and building our ultimate attacker and then our other agents that help customers secure their environment, uh, and make themselves better prepared for the AI wave.
Ashish Rajan: Awesome. I mean, both of you have black hair, which is, and it's like after spending so much time in red teaming, I'm just surprised. But it, it's a good thing. It makes me envious of you guys. One thing that has been top of mind for people who, at least most practitioners, and we are at BlackHat this morning, a lot of practitioners obviously have been used to the idea of a manual pentest, whether it's once a year, once a quarter, however people do it, whether it's internal, external, many versions of it.
Ashish Rajan: Obviously, you guys are, have gone down the path of using agents to do a lot of the testing. I'm curious in terms of why do you s- a feel is right th- this is the right time to start that conversation and how different is this to Uh, me thinking that I can automate, uh, my script and pen test.
Ashish Rajan: Uh, I'm curious 'cause you mentioned agents, so maybe we can start there, and then we'll come back to you, Evan. [00:04:00]
Travis Lanham: Yeah, I, I think, you know, the big thing that's happened, right, is that AI has unleashed this new wave of attacks that have a scale, speed, and sophistication that we haven't seen before, right? Uh, and at the same time, we've seen this democratization of advanced offensive capability, right?
Travis Lanham: You know, when we had the EternalBlue moment, it was a moment where now, you know, every teenager in their bedroom had access to nation-state grade offensive capability and malware, right? Mm-hmm. And we're seeing the same thing with AI today. It's enabling everyone to be much more advanced. It's creating a lot of disruptive scenarios, and it's even created the, you know, condition of having an autonomous agent that can go and do things that weren't intended.
Travis Lanham: And so in that backdrop, every organization now needs to be preparing themselves for this world where the threat landscape is dominated by AI attacks. We call these hyperattacks, where you have a swarm of agents coming at an organization, testing every part of their perimeter, testing every part of their [00:05:00] external, testing every part of their organization relentlessly.
Travis Lanham: And so, you know, we, you know, in the past would talk about the idea of, quarterly or annual pen tests. Then we talked about the idea of, like, continuous pen testing. But what really matters today is being able to withstand these relentless AI attacks
Evan Peña: yeah. The truth is we've evolved into a place where AI can now augment human workflows, and we've seen it in multiple industries, not just so- cybersecurity, right?
Ashish Rajan: Yeah.
Evan Peña: And so having come from an era of having to do this manually and having true tradecraft and being able to train agents to do things the way that we would, and you see the same quality outcome, but now you can scale dramatically more, like, why wouldn't you do that, right? Yeah. And so, uh, it's been a fun journey, uh, building these agents with Travis and, and teaching them how to do it the way that we would do it.
Evan Peña: 'Cause, like, an out-of-the-box model doesn't quite do it exactly how we would do it. And so, like, having true tradecraft and being able to, like, train these agents the way we would do it, you can scale so much more. You now have three things that are no longer a barrier. You have [00:06:00] significant more coverage.
Evan Peña: Yeah. Whereas before, you couldn't cover the entire attack service. You have any sort of one A enterprise, and you have a, a finite amount of team, human capital- Yeah ... behind that to try to, to solve that problem. You just... There's not enough time. Yeah. There's not enough, uh, expertise, and there's not enough coverage.
Evan Peña: So now with AI agents Evan has a swarm of a million, team, uh, and so versus, like, maybe 200 people. And so that this, this scale is just so much more than ever before, and that's... It's cool to see how much more coverage that we've, we've been able to cover than ever before. And-
Travis Lanham: And, and we, we, we talk about this kind of idea of, like, planet-scale intelligence with AI meeting elite tradecraft, right?
Travis Lanham: Mm. And what these new attacks of the future look like, right? Yeah, yeah. You know, we just did one of these hyper attacks to help an organization where we had, you know, 13,000 concurrent attacks, 26,000 agents, 40 billion tokens. You know, this is swarming across their entire organization with 25,000 exposed services, right?
Travis Lanham: It's just a massive scale of attack that wouldn't have been possible previously. Yeah,
Ashish Rajan: [00:07:00] yeah.
Travis Lanham: Right? And so being able to really pressure yourself against these relentless AI attacks is something every organization needs to do right now.
Evan Peña: Actually, you mentioned something earlier that, that resonates well to that point is, you know, what's the difference between my script doing it and- Yeah, yeah
Evan Peña: an AI, right?
Ashish Rajan: I was gonna ask that 'cause I imagine all the purists listening or watching this will be like, "Come on, Evan." Oh, I I have... And I think I probably would add to this, not just a script, but How different would that be, me as a pen tester adding a Claude Code, may building some skills internally, and to your point, having multiple Ashish behind me?
Ashish Rajan: 'Cause I imagine it feels the right thing to do. How far or how close is that to the reality of being this, the scale that you guys have built? And when you see other people kind of think that, "Oh, I, I can do this myself, man. Don't... Thank you so much." Like, so- Yeah ... what do you guys think about that?
Evan Peña: I- I'll address the script one first- Yeah
Evan Peña: and then I'll let Travis talk about, like, the difference between a skill and, like, what we've built, in that the, the scripting is like, think of, like, SOAR for pen testing, orchestration, automation. It's very narrow focused. It's [00:08:00] like if this one situation happens, I can account for the script. Yeah. Or a Nuclei template, for example, that's specific to one scenario.
Evan Peña: It doesn't have the autonomy to think around the scenario or be flexible or determine if I, if I... Like to say one of the scripts is determining if there's a SQL injection vulnerability.
Travis Lanham: Yeah.
Evan Peña: What happens after the SQL injection? What happens from there, right? Like, it's cool about AI agents, like, it, it can actually continue the kill chain, continue the path.
Evan Peña: Or if there's something like a compensating control in front of that SQL injection, it can think around it and still compromise the SQL injection as a basic example of how it can think on its feet. It can adapt to the situation. It can think exactly how we think, 'cause we don't just train these agents to do it the way that we would from a methodology perspective.
Evan Peña: We literally think, we sh- we teach it how to think like we do, and, like, that's a big differentiator as well.
Travis Lanham: I, I think just to add to that, right, at Google we actually had the market leading product for breach and attack simulation, which is the scripting approach that, you know, Evan was talking about, right?
Travis Lanham: And it was just never flexible, [00:09:00] adaptable, or evolvable like- Yeah ... a real attacker was. Yeah, exactly what Evan said. And so it was providing this false sense of security or these very narrow use cases where it's, that's, you know, one, one-millionth of what you really need to be doing, right? Uh, of what you need to be preparing yourself for.
Travis Lanham: And so then to your other question, right, when you think, you know, when we think about, what we're building it's really this idea of, what you have this idea of like having the skill or, you know, having like what you want an agent to do, right? But so much now has become about the system of record that those agents become.
Travis Lanham: In security, we've never really had a system of record. We've had configuration management databases, we've had SIEMs, you know, we've had like all these different things that try to put together this picture of what's the ground truth in the environment, right? What's the Google Street View for what's real in the environment?
Travis Lanham: Mm. And that's what we built, where we have our agents fan out over an organization, over a network, build up this real-time knowledge graph of everything they see, index that, and then be able [00:10:00] to attack it relentlessly with AI. And so that gives us, kind of scale that can't be matched. That gives us speed with optimized agents, optimized serving infrastructure.
Travis Lanham: You know, we've, you know, done a lot of post-training work to make that, you know, go really fast. Uh, and then sophistication through, you know, expertise and judgment really, soaked into the agents. And then something that's, you know, top of mind for everyone right now is safety, right? From day one, we built a safety infrastructure that provides control around AI.
Travis Lanham: That's both deterministic control and then a safety agent that we've trained with world-class red teaming expertise over a long period of time in a lot of different diverse environments, so that we can build up, kind of a judge or a critic that can evaluate, "Hey, is this thing the agent wants to do safe or not safe?"
Travis Lanham: So that, you know, it can't escape the cage.
Ashish Rajan: Is that what ties into, I think to what you were saying as well, Evan, where, uh, a lot of the questions that people have is that, "Hey, uh, the reason I used to go for a bug [00:11:00] bounty program or a pen test was because it was business logic flaws that I could not normally get."
Ashish Rajan: And o- obviously to what you said, the whole-- That came from the idea that I, it, I could not chain multiple vulnerabilities together. When it w-- Whether it was low sev or high sev, whatever, people were able to build a kill chain around it. One of the things that people talk about automated pen testing as a concept is that h- and maybe it ties back to your safety question as well, how does it know to stop?
Ashish Rajan: Like, you know how usually when you do a SQL injection, you realize you, okay, I've taken three steps, and now I know I'm at prod data now. Usually, you'll stop. You go, "Okay, at this point in time, if I do anything, this is probably b- beyond the scope of the contract." Is that easy to do with automated pen testing in terms of, obviously, because you guys have worked with multiple customers, h- how do you balance that approach?
Travis Lanham: Yeah, I, I think, you know, there's kind of the piece on, like, having a halting criteria for the agent, right? Or a stop criteria on, you know, that we've achieved the objective that we set out to achieve. Right.
Ashish Rajan: Right.
Travis Lanham: But what that really gets us to is this baseline [00:12:00] verifiable state that we want the environment to have, right?
Travis Lanham: You, you can think of it as in so many different domains, we have this idea where we wanna test something, right? Mm-hmm. Like in software engineering, you write some code, then you test the code to make sure it does what you say. In security, we've had to, you know, to your point at the beginning, right, that test part has been once a year.
Travis Lanham: Yeah, yeah. That test part has been only a tiny fraction. Now, with AI, we have the ability for... in software engineering, if you tested, like, 1% of your environment, and you only tested once a year, that would be crazy, right? Mm. You know, you... If you have, you know, a good software product, right- Mm-hmm ... you know, or a good software system, it has, you know, like, 99 or 100% test coverage, and those tests are executed every time you submit the code, right?
Travis Lanham: In security, we've had, like, the inverse of that. Yeah. Right? But now, for the first time, we have the ability to have this verifiable baseline that we can just continue testing your environment against, and do this in a way that scales well, that's understanding the risk in the environment, that's able to go and apply that expert human judgment [00:13:00] at scale, so that as the environment changes, and they're changing faster than ever with AI accelerating the pace of digital change, there needs to be that baseline that everything is verified against.
Travis Lanham: And to kinda, like, dive a bit
Evan Peña: deeper into the architecture of how, uh, Travis and team have built our agents is we have what we call campaigns, which you can set up a campaign that has multiple attacks. Yeah. And let's just say, for example, one of the attacks is, "Hey, compromise this web s- web application."
Evan Peña: Yeah. Basic example. There you go. And then you provide context. So that's rules of engagement, its objectives, and so that context gets sent to this Overwatch agent. Mm. And then the Overwatch agent spawn... Essentially, it spreads out, right? And it starts, like, spawning scout agents, recon agents, and then we have what we call hunter agents, and there's a specific hunter agent for SQL injection, and there's, like, hundreds of them for different attacks that it's specifically trained for.
Evan Peña: And so the SQL inject- hunter gets spawned-
Travis Lanham: Yeah ...
Evan Peña: and then it goes down the SQL injec- injection path, fully compromises the SQL injection, [00:14:00] and then it will send the data back to the Overwatch. Like, "Hey, I got this database credential, or I got this." Right. And then the Overwatch takes the context and- Yeah
Evan Peña: determines where do I go from there. But the Overwatch agent... I mean, the hunter agent will stop because its only objective is the SQL injection. Yeah. And then the Overwatch agent determines, like, what's the larger objective here for the ex- uh, you know, the type of scope for the campaign and the attack.
Ashish Rajan: How do you guys manage hallucination then? 'Cause I g- and the reason I say that is because, obviously a number of years you've spent in AI, the, and you guys spent Google, Mandiant, and a lot of people are still... And it is still a thing. Hallucination is still a thing. How do you guys balance that? 'Cause a lot of people, the reason they won't go for a bug bounty or a pen test is because the, it's A, it's a third party, B, it's like the Well, let's just say there's an assumed trust when a human does it, even though the human may miss the exact same things- Yeah
Ashish Rajan: but there's an assumed trust there, right? With this particular ecosystem, how do you maintain or balance the hallucination for, to your point, you have like a specialized agent for SQL injection, specialized for cross-site scripting, all... How do you balance that? [00:15:00]
Travis Lanham: So we, we've actually also built and trained our own verification agent and system where it's not...
Travis Lanham: Yeah, and that includes both deterministic pieces as well as non-deterministic pieces, where we'll go and take, "Here's a kill chain that we found in the environment," and then be able to go and actually provide the proof that was in fact a real kill chain, right? Right. And traverse those same set of steps and be able to reproduce it, right?
Travis Lanham: Because to your point, right, it's like, hey, if you tried to test something and you can't reproduce the test, you know, it's flaky or it doesn't really make sense. Yeah, yeah. It's like going to math class and saying, "Show your work," right? Yeah, yeah, yeah.
Evan Peña: That's literally like the, the idea,
Ashish Rajan: right. And, and this goes back to, building trust in the output as well, right?
Travis Lanham: I think you hit on a, a piece that's really important here, right, which is this idea of a third party- Right ... right? And having that third party seal of trust, right? I think that's a really important thing that we found with, you know, the organizations that we work with, that they value a lot, right, on basically being able to use our system, but then also have this layer of third party approval.
Travis Lanham: Yeah. Right? Because, you know, you to your point about a skill file, right, like a skill file isn't an audit, right? A skill file [00:16:00] isn't proof that, you know, you did what you said or the agent did what it said, right? Yeah. I think there's now very much this, you know, belief when, and kind of understanding from people who interact with real AI systems that there needs to be kind of that level of proof.
Travis Lanham: There needs to be that level of this actually did what we said, right? Or, you know, this was... If it wasn't something that could just be verified by like easily looking at a screen or a dashboard, you know, what actually happened.
Ashish Rajan: Yeah. Yeah. I mean, skills just means I'm cool. That's basically what- I figured this out, guys.
Ashish Rajan: You got skills. Yeah, yeah, yeah. I, it's, I, I mean, it's a street cred kind of a thing as well, I guess, at this point in time, 'cause all the bug bounty people have started talking about what kind of skills they have and stuff. But- Mm-hmm ... you mentioned earlier about the, the coverage of it as well, right? I think at the moment, uh, you mentioned the countries monitoring part, and w- likely, uh, most likely even now, people are still doing once a year or twice a year kind of schedule because that's the regulated requirement.
Ashish Rajan: Why is there an em- a sense of urgency based on what you guys are [00:17:00] seeing for having a continuous coverage and coverage across everything that is interfacing at least?
Evan Peña: Yeah. I mean, I think what from my observation, like having done this for so long-
Ashish Rajan: Yeah ...
Evan Peña: we just didn't, we weren't able to have enough time to like cover everything.
Evan Peña: So I think things are surfacing now that have never been seen before. Like, "Oh, I didn't know that asset existed externally," or, "I didn't... Oh man, that application's still around?" Like- Oh ... I had no idea. Because generally we take the path of least resistance to accomplish an objective- Of course, yeah ... when it's a human-led assessment.
Ashish Rajan: Yeah, yeah. You have two weeks to prove yourself.
Evan Peña: Exactly. Like how are you gonna cover the full attack surface in that amount of time? Yeah. How can you discover every kill chain that exists- Yeah ... with a super finite amount of time? So like now it's just unprecedented times of like the amount of coverage and scale that we have with the AI agents, and so we're just surfacing things that have just been around for a while that just people do not know about.
Evan Peña: Mm. And we are finding very critical, high... Like, I'm talking like, when I say critical, I'm not talking about like a stored cross-site scripting. I'm talking like RCE on an s- internet-facing server that got us access to a cloud tenant, a Kubernetes cluster- Wow ... uh, breaking [00:18:00] out from nodes to clusters and accessing other tenants, or even getting to a corporate environment.
Evan Peña: Like it's big impact things. Like you don't know what you're gonna get. Yeah. So we just can cover so much more, and then people just weren't able to see some of the stuff that we're surfacing.
Travis Lanham: A- and I, I think we've now seen this repeatedly, right? You know, when we talked about that hyper attack that we did, right? We found, you know, real issues in this environment, uh, that the organization went and addressed right away as an incident, right? You know, there's this idea that if you know, in yesteryear pen tested a few of your apps, if you pen tested some of your network, if you, you know, did some form of security testing and assessment, that it was good enough, right?
Travis Lanham: Mm. And now with AI, that's just no longer the case, right? Yeah. You know, the, you know, the AI wave is coming for the cybersecurity environment, right? We've now seen this over, you know, 2026 has been the year of AI is now here in your cybersecurity environment, right? Yeah. It matters, right? It is now, you know, across, you know, all the organizations we worked with, you know, defense, retail, healthcare, finance, you know, [00:19:00] critical energy infrastructure, right?
Travis Lanham: We've always been able to find way more than the organization thought existed, as Evan said, right? It's this idea that if you now have this swarm of relentless agents just pounding against your organization, that is a far different threat landscape than existed even 12 months ago.
Ashish Rajan: Is there a question of building You know how most of the audit, they go down the path of a human audit, human done pen test versus an automated pen test.
Ashish Rajan: Is there, uh, questions from customers around, "Hey, how do I build some..." Uh, and the-- it kind of extends in that trust example that I was asking about earlier, where normally people go, "Hey, I'm gonna go to, I don't know, like just say a Big Four."
Ashish Rajan: I can stamp the name of KPMG or whatever. People are op- finally opening up to the idea of automated pen test, because a lot of people re- the reason why we had that two-week window is because I was too scared to find out stuff that I don't want to know about.
Ashish Rajan: So I give you a very short window that I can go back with a rubber stamp saying, "Hey, someone pen tested this." And to what you said about the hyperscaler attacks, suddenly how comfortable do you [00:20:00] find people are? Like, are they opening up to the idea and realizing that, oh, actually, you know what? To what you guys said, this is here to stay.
Ashish Rajan: And maybe Hugging Face is a great example as well of, uh, this, which I think is you guys had some thoughts on that as well. So I'd love to unpack that. But is people, are people opening up to the idea of automated pen testing a lot more, or is it still like you see, blips of radar here and there that the forward thinking are the only ones doing it?
Ashish Rajan: I'm curious.
Travis Lanham: No, I, I think everyone is, is- Right ... uh, I think everyone is now open to the idea that the attacks of the future are AI. Yeah. And that's what people need to be preparing themselves against. I think we're also going to look back on this, and the idea of pen testing is just kind of going to, like, be a, like, legacy kind of concept.
Travis Lanham: It'll be a skill.
Ashish Rajan: Yeah. Yeah.
Travis Lanham: Well, I, I think it's not even that it'll be a skill, right? I think the reality is, is that AI attacking you looks far different than, like, what pen testing was, right? And now that AI is cyberattack- Yeah ... right? Like, AI is the offensive actor in the cyber domain. Yeah,
Evan Peña: [00:21:00] yeah.
Travis Lanham: That is what matters.
Travis Lanham: Yeah. And so if you aren't preparing yourself for those relentless AI attacks, like, it's not about, like, you know, making pen testing more efficient, right? It's not, you know... It's, it's about, like, taking the kernel of the truth that was there, right? Mm-hmm. Which was this idea of verification, right? Yeah. And being able to understand what does verification look like in an AI world.
Ashish Rajan: Yeah. Also-
Evan Peña: I don't
Travis Lanham: know why someone would do a
Evan Peña: traditional pen test at this point, because when we just do, like, a proof of value-
Ashish Rajan: Yeah ...
Evan Peña: we'll find more in that one 48 hours than people will find in an entire external pen test done with a two-week, you know, two-person kind of team.
Ashish Rajan: Yeah.
Evan Peña: And it's significantly more because we're just covering so much more.
Evan Peña: So I, I tend- I almost think, like, not to be too harsh, but it's almost negligent to, like, not, you know, to do a traditional pen test with the amount of scaling capability that not just we have, but, like, what real nation state threat actors have. I mean, they have the same capability that we're building, I'm sure.
Evan Peña: And so who's to say they're not gonna go after the same infrastructure that we're going after and gonna find [00:22:00] the same things and extort a customer for whatever it is that, you know, they would want to extort them for? Yeah. So I think it's a, it's, it's such a critical shift in the entire industry, as I'm sure you're seeing-
Ashish Rajan: Yeah
Evan Peña: Of, like, this is what you need to be doing to truly understand if you're prepared for that or not.
Ashish Rajan: Yeah.
Travis Lanham: And we, we've already seen this start to happen, right? The new attacks that are shutting down water infrastructure that are, you know, impacting organizations today, they're AI powered. The, the threat actor is AI.
Ashish Rajan: Yeah. How do you see the difference, though? I think a l- And I'm curious, what did you guys learn from the Hugging Face thing that you're able to ki- And it kind of cements on what both of you are saying as well, that the agent attacks are here. So it's no longer that it's happening privately.
Ashish Rajan: It's public information as well.
Ashish Rajan: I think Hugging Face is, was just the more recent one.
Evan Peña: Yeah.
Ashish Rajan: What were some of the learnings over there that people who probably are the ones who do the pen test every two weeks or so, but in a two-week window what are some of the learnings that you guys had there that made you guys even go, "This, this is exactly the reason why people need to be doing more of the [00:23:00] automated pen test versus just a manual two-week once a year thing"?
Travis Lanham: Yeah, I think it brought to light this idea that AI has this capability-
Ashish Rajan: Yeah ...
Travis Lanham: that everyone who has access to those AI capabilities can do this, and the number of people who have access to those AI capabilities is huge, right? Mm-hmm. Especially, you know, with, you know, the open-weight models, kind of the diffusion of capabilities, there's now this capability that as everyone as a baseline has- Yeah
Travis Lanham: that is highly advanced, right? You know, that three years ago would've been considered only the purview of top-tier nation-state threat actors, and now that's available to everyone. So I think there's really this, awakening that's happened for folks in the security community, uh, you know, that this is something that's here today that matters, that needs to be addressed now, and that really calls to light the importance of safety and governance around these systems where, you know, you can't just have a skill file that you pop into your, your Claude Code Harness, right?
Travis Lanham: That's what happened in the incident, right? You need to actually have something robust that's a [00:24:00] system that's built to ensure safety, that's a model that's trained specifically for that judgment, where you're building something that has that expert discretion that's not just kind of this mixture of everything that was on the internet.
Travis Lanham: There's a lot of bad stuff that's on the internet. Yeah.
Ashish Rajan: Well, what about the whole Mythos conversation as well, where a lot of people, unfortunately or fortunately, are waiting for Mythos to be avail- become available so they can become, they can become super skilled, for lack of a better word. Does it-
Travis Lanham: It is, it is available, right?
Travis Lanham: I mean, yeah,
Ashish Rajan: like- Oh, to the limited company, but I think more in the context of the idea that, "Hey, oh, maybe my skill is not good enough at the moment- Mm-hmm ... but if I had access to Mythos or I don't know whatever the cyber model that GPT has-"
Travis Lanham: Or Qimi K3.
Ashish Rajan: Yeah, yeah, yeah, or Qimi or
Travis Lanham: GLM. This level of capability is- Yeah
Travis Lanham: now widely available. I think that's something that, has, like- kind of, soaked in or percolated in the ecosystem, but I think has not really, like, had the fine line to it that this level, you know, in April, this capability was highly reserved, right? Yeah. Oh, yeah. A few [00:25:00] people had access to it.
Travis Lanham: Today, this capability is available to anyone who can run Kimi
Evan Peña: K3. Yeah, I would argue, like, Kimi K3 and even GLM, like 5.2 or the, the latest GLM models are even more capable than Mythos. And, and the reason why is 'cause, like, we tested Mythos with a customer of ours that was part of Glasswing.
Travis Lanham: Mm-hmm.
Evan Peña: And, like, we got a bunch of refusals.
Evan Peña: We had to keep... Like, we had to prompt and, like, engineer our way into getting it to work, and then we compared it to the same results of what we would find against the same targets, and we find more actually than what Mythos found. Not because of the model itself, just it's because, one, we got significantly less refusals with other models, and the other models are very capable.
Evan Peña: Um, and I think Mythos was just a signal to the industry, and, like, because it was not publicly released, I think people got scared of the situation. Yeah. But at the end of the day, like, the model that Travis just mentioned with Kimi or GLM, like both of those are very capable models and very violent in nature in terms of, like, they don't refuse anything that you wanna give it.
Evan Peña: And so the, the [00:26:00] main thing that people need to think about is, like, not who has access to Mythos and what they can do, but, like, anyone has access to these other models. And our nation state threat actors or our any sort of bad threat actor is gonna be using these models 'cause they're significantly cheaper, they're easier to work with, and they're also very capable from an expertise perspective.
Evan Peña: Like, the, the trainings on those is pretty good. And you- And you can post-train it to do just like what we're doing. Have an expert team of hackers train these models post-train, and they're gonna be even more capable.
Ashish Rajan: Also to, to what both of you are saying I, I guess the bigger gap to be filled over here, if someone was to think of, "Hey, I'm gonna build a skill or a super skill or whatever," it's just not the fact that I have built a, a series of SQL injection skills, cross-site scripting skills, whatever, insert RCE, everything, just add, the list goes on.
Ashish Rajan: It's more also the consistent understanding of the newer environments that you guys are bringing in and the shift in models as well. The more capable the model, and doesn't really matter if an open model, closed model, whatever are you [00:27:00] finding that The people who are here at the moment in, in BlackHat trying to build a, hey, I wanna have an agentic security program, or I wanna build an AI security program or uplift an existing one, right?
Ashish Rajan: What's a what's a good place to wet your feet in getting comfortable with the whole automated pen testing thing? 'Cause I, I think a lot of people already believe that, you know, and I almost feel like when I talk about I simplify the example of calling it two weeks and two people, whatever, be the team, I guess, based on the budget.
Ashish Rajan: But there are people who have internal pen testers as well. Uh, large banks, all of them are internal pen testers. Uh, some of them have access to Mythos and perhaps the engineering capability as well. Where do you find is the human loop in this equation where... 'Cause I feel the more, like you said, you build in agents, you have the expertise.
Ashish Rajan: Is the human in the loop supposed to be on your side, or is it on the, the customer side? Where does that human in the loop exist, and should it still, is there a need to still have that?
Travis Lanham: Yeah, it can be on either side, right? Yeah. So we, offer our platform for customers who wanna [00:28:00] operate it with their own operators, and then we also, you know, provide for organizations that might not have that level of expertise, you know, our own, you know, third-party seal of approval.
Travis Lanham: And we've actually seen a lot of organizations that want both, right? Mm-hmm. Where they want to, you know, they have a deep understanding of their internal network, of their internal setup, but then everyone wants that third-party seal of approval, right? That has to exist.
Ashish Rajan: Yeah.
Travis Lanham: That is something that's fundamental to security, right?
Travis Lanham: You know, you don't trust your finance team to audit your own books, right?
Ashish Rajan: I mean, that'd be great if that was allowed. Yeah,
Travis Lanham: it's, it's-
Ashish Rajan: I just want the extra zero in the end.
Travis Lanham: Yeah, exact- it's not allowed for obvious reasons. Yeah, yeah. But, you know, wh- when that's happened in the past, it hasn't gone very well, right?
Travis Lanham: Uh, and so, you know, I, I think what we see, though, really is this idea that, you know, in 2011, right, there were people who thought that they could build their own what now has become EDR, right? Yeah. I don't think there's anyone today who thinks that they can, you know, go and build, a system like, you know, CrowdStrike, where you have this massive scale, you have, you know, this expertise that's constantly training things to get [00:29:00] better and better.
Travis Lanham: You have visibility into so many different environments and so much diversity. Uh, and you know, the- there, you know, a great technical partnership that we've done on being able to say, "We found a kill chain in your environment, and now we can train your CrowdStrike to get smarter to block that," uh, and be able to, you know, do that across different control surfaces, on the endpoint, on the network, on identity.
Travis Lanham: Being able to really get into this loop of we have, you know, this red AI training blue AI to better secure your environment in an autonomous way. And then to the human in the loop part, right, I think it's, it just depends on different things, right? You know, I think there's parts now where we've had organizations get very comfortable that in certain scopes with certain sets of controls, that they're okay with, you know, the red agent going and being fully autonomous.
Travis Lanham: Uh, you know, there we're kind of, at a different stage with autonomous remediation, right? Again, in some scopes and some actions, uh, you know, organizations are comfortable with those steps being done autonomously, and in others, they want a human to review to make sure that, hey, this isn't taking, uh, you [00:30:00] know, down a critical, critical, you know, web server.
Travis Lanham: This isn't taking down a critical, production network system. This isn't taking down a critical service, uh, or doing something that would disrupt the environment. Because a lot of organizations know that they have a lot to go on just kind of the generic you know, business digital QA, right? On like how do I go and test, a, you know, retail shopping site to make sure that the checkout flow still works, right?
Travis Lanham: If I don't have great test coverage for that, then if I'm going to making automatic ch- automatic changes in the IT environment- Mm-hmm ... that could be highly disruptive to the business, right? And so, you know, it's still worth having a human in the loop for some of those things to make sure that the right thing's happening.
Travis Lanham: But I think every organization understands that these things are becoming faster and faster. The machine speed offense begets machine speed defense, that everyone needs to be moving much faster than they were before, right? You know, we talked about the two weeks, we talked about the one year, right? We talked about all these things that were kind of at human speed and human scale that were possible in yesteryear.
Travis Lanham: [00:31:00] But today, we're at machine speed, machine scale, machine sophistication.
Ashish Rajan: Curious also, uh, for people who may or not have access to the thing that you guys are doing, what... And probably Hugging Face is a good example to kind of bring up again. What does it look like when an agent is-- Like, you guys have the red team agent, the blue team agent.
Ashish Rajan: For the blue teamers who are listening or watching this going, "I don't have access to that at the moment," but how do I even know that something, quote-unquote, "agentic" is in my, in my environment? And is it the fact that, I don't know, the time speeds is very low, which used to be a pattern some people used to have for au- differentiating an automated action versus a human action?
Travis Lanham: Yep.
Ashish Rajan: Are there things that... Uh, and you guys are obviously seeing both sides. You're seeing the blue team side, the red team side. What's something that you have learned that people can use as an indicator for, "Hey, that's not a human"? That's something, uh, let's just say agentic if that word is
Travis Lanham: Well, I, I think that's kind of the key to our approach, right?
Travis Lanham: Is that you need this relentless swarm to come at you. Yeah. Because we can show you exactly where it went, [00:32:00] right? We can show you step by step the audit trail of everything that these agents did in your environment, right? You have that unified, offense and defense visibility together in one.
Travis Lanham: Yeah. Because it's not a real threat actor who's in your environment, where they're the only ones who know where they've been, what they've done, how that was effective, right? Like we have a whole detection response paradigm that takes exabytes of data and tries to process them every year to go and like f- stitch together these kill chains from the observer perspective or from the defender perspective only.
Travis Lanham: And I think that's just a fundamentally flawed approach in this new AI-powered world, right? You know, with a bunch of these, you know, hyperattacks where we've gone and helped organizations prepare themselves against this kind of relentless, continuous swarm-
Ashish Rajan: Yeah ...
Travis Lanham: uh, you know, what we've found is that oftentimes these aren't, you know, these actions aren't detected.
Travis Lanham: The agents aren't detected in the environment. Uh, you know, it's not even that they're not detected, but it's very hard to just see the activity at all, right? Mm. Because there's lots of gaps, right? If you think about an environment previously, it's like, [00:33:00] okay, we have like one camera at the front of the 7-Eleven, right?
Travis Lanham: But what about when you need to like make sure that you're capturing everything, right? You know, if you're a bank, you need to make sure that you've got cameras everywhere. You need to make sure that you're seeing everything, and the first step is going and having, the red go through and show you, all right, here are the blind spots, right?
Travis Lanham: We've married up that with the blue perspective or with the obser- you know, the defender's observability perspective and be able to show, here are, you know, the P0s that you have for improving your defensive visibility. Uh, and I, I think that's kind of a huge piece that organizations are now coming around to, that it's not even about like, oh, are there signatures that I can look for?
Travis Lanham: Are there, you know, patterns that I can look for? But am I e- even looking at all, right? Yeah. Or, you know, how do I make sure that we have good coverage?
Ashish Rajan: I, I love this also because, uh, you mentioned the detection response thing as well. It's a, it's an interesting conversation, right? If, if we start using autonomous testing for blue team, red team, and, uh, using AI agent capability with experts in the background providing [00:34:00] that what does that mean for the security team moving forward?
Ashish Rajan: Like if we were to fast-forward this, obviously right now a lot of people are mixed at the moment. There's, there's this, uh, even adoption-wise for AI, there's a mixed coverage as well. Uh, say a couple of years down the line, where do you see this go in terms of how does the security ecosystem evolve?
Ashish Rajan: Do you see the detection response thing kind of shape into what you guys are doing, or does it kind of split away, or does it even disappear for that matter as well? Yeah. I'm curious, what do you guys see in the work you guys are doing?
Travis Lanham: Yeah, I think there's this like fundamental change to it, right? When we say detection today, we often mean, you know, control point detection, or we often mean putting data into a SIEM and doing detection there.
Travis Lanham: But the new form of detection is actually just doing the attack, right? The attack is now easily scalable, at high speed and at the sophistication of the real attacker. And so the way to detect is now to just attack yourself, right? Yeah. And to see what happened. And so yeah, I think there's this new part of the detection [00:35:00] paradigm, and then the response paradigm is you now have better fidelity or better data for detection than you ever have before.
Travis Lanham: Because you're not just reacting based on these kind of lossy sensors and things that you're trying to capture what's happening in the environment. You actually see the path, right? You know exactly what happened, and so you can go and respond much better. So I think, you know, to your point about like what is the security team of the future going to look like, I think it's this idea that a lot of the traditional work that's gone into detection and response- Yeah
Travis Lanham: is going to really evolve to, okay, the new detection signal is what red can do, right? And now I need to go and analyze that. I need to go and be able to adjust, you know, the like, you know, see what the response plan is, see what blue says, and then approve that. And then, you know, just keep, you know, gain the environment into better and better shape.
Travis Lanham: And so it kind of like flips on its head this paradigm that we've had of you wait for the alert to come in, you respond to the alert. Now you wait for the kill chain to come in and you say, "Okay, here's how we're going to fix this. [00:36:00] Here's how that, corresponds to the long-term governance changes we're making our en- in our environment, the long-term hardening plans that we're doing in our environment."
Travis Lanham: And so it now becomes that the job of the security team is to have that discretion between what's that tactical, put the tourniquet on, contain the, you know, the infection, and have the compensating control in place or the mitigation in place. And then where does this fall into the long-term plan of getting the whole environment to a better state?
Ashish Rajan: Would we... going back to the security programs people are building today, and to your point, who are planning for, hey obviously I'm not planning for the next six months, probably planning for the next couple of years for what my AI security program would look like. Is the understanding there then for, like say the evolving world of protection to what you said, then you d- instead of waiting for something to happen, being proactive about it and just going, "Hey, let's just go on the offense," rather than trying to just sit on and wait for hopefully Hail Mary, nothing happens.
Ashish Rajan: What do you think is a good place to start having this kind of injection, for lack [00:37:00] of a better word? 'Cause a lot of people obviously have, Let's just say they obviously have pen testers, they have GRCs. So all these processes defined, where do you find your customers have the quick wins to show ROI to the organization going, "Oh, okay."
Ashish Rajan: 'Cause there's obviously a skepticism for should we still do the KPMG or whatever the Big Four thing is to validate, hey, this is the value. Like, what are you finding as ways people are able to kind of put you guys forward as, hey, this is more valuable than us spending our time, or which I imagine it happens over time.
Ashish Rajan: It doesn't happen, like, on a flick of a second in terms of, well, I'm
Travis Lanham: sure it happens in- For, for a lot, yeah, for a lot of organizations it does. It does. When you've experienced a hyperattack, it provides that clarity, right? It is a, you know, it's not just a, oh, here's, you know, a PDF report- Mm-hmm ... kPMG or someone else would produce, right?
Travis Lanham: And there, there's value in that, right? But it's really this idea of, oh, this is an incident.
Ashish Rajan: Yeah.
Travis Lanham: I need to treat and respond to this like I would an actual breach.
Ashish Rajan: Yeah.
Travis Lanham: And it provides that clarity or that crystallization of here's what a relentless attack, here's what a hyperattack looks like in my environment.
Travis Lanham: That's [00:38:00] what I need to go and prepare myself for. That's how I'm going to take my SOC and operationalize them as, you know, now this, like, vulnerability operations center, uh, or this, you know, exploit, you know, operations center- Yeah ... on going and fixing those issues.
Ashish Rajan: Yeah. And so, so would you say the easy place to start or, or not an easy place, but at least the, uh, the quick win for conversations where it is not that one-second thing and they're trying to position you guys in from a, "Hey, I wanna do more automated, but at the moment I have to balance the two," where do you find...
Ashish Rajan: Is it more the web apps? Is it more the cloud apps? What's the-
Travis Lanham: The hyperattack. Ex- The hyperattack. Exactly. No, it is, it's the hyperattack. It's- Certainly. Yeah, every- Have us go at you,
Ashish Rajan: yeah. Yeah. I mean, technically, to your point, it is the reality. You can say, "Oh, I, let's go for the easy one." You're like, actually, technically, I mean, the hard one is sounding like the more possible thing in the future right now.
Travis Lanham: Yeah. The, uh, I mean, it's like something that wouldn't have been possible before is possible today. Yeah. Right? And every organization needs to do this, right? Every organization needs to [00:39:00] have this understanding of what does it look like to have this relentless AI wave safely- Yeah ... come against your organization, right?
Travis Lanham: And that's no lo- it's no longer about web apps, it's no longer about services, it's no longer about the network. It's about this totality of what does it look like to get hit by a hyperattack, right? Mm. And then to be continuously doing that, so you, you hyperattack yourself, you go and fix things, Blue Talks and, you know, tightens the gap, right?
Travis Lanham: And then you need to keep doing it, right? Because you have more change going on in your environment. So it really is this kind of now is the moment that every organization needs to understand what does it look like to have that relentless swarm coming after my organization.
Ashish Rajan: Awesome. No, thank you so much for that.
Ashish Rajan: That was all the technical questions I had. I'm doing this section called the you laugh, you lose.
Travis Lanham: Yeah.
Ashish Rajan: Do you guys have jokes? The way you- You, you promised that you were gonna- I promised- Start with the jokes ... I'll joke, but I promised you guys have a joke as well, apparently. I'll tell you my joke if I- Okay. Well, let's just do it one-sided then. So the whole idea is, uh, I've got a couple of jokes, so, because you do have two of you. The whole idea is you get [00:40:00] five seconds to react to the joke.
Ashish Rajan: And if you laugh, you lose. That's pretty much the, the, the play.
Ashish Rajan: All right. Here's my first one. Why was the SQL injection so good at poker?
Evan Peña: Why?
Ashish Rajan: It always knew when to union.
Evan Peña: Ugh.
Ashish Rajan: Nice.
Evan Peña: That's good.
Ashish Rajan: Uh, I've got another one. Uh, why are bug bounty hunters terrible at dating?
Evan Peña: Why?
Ashish Rajan: They keep looking for red flags.
Evan Peña: Fair. Oh.
Ashish Rajan: Uh, what wait, is it good enough?
Ashish Rajan: Is it-
Evan Peña: It's good. You guys just think that it's- I got, I got one. You inspired me.
Ashish Rajan: All right, okay.
Evan Peña: This is literally on the fly.
Ashish Rajan: Love it.
Evan Peña: Okay, okay. Uh, what do a hacker and a tree have in common?
Ashish Rajan: What?
Evan Peña: They've both got root.
Ashish Rajan: All right, okay. No, I, I, I like that one actually. I... Well, it is, I mean, so to what I was telling you guys earlier, that, uh, it's the...
Ashish Rajan: It's so hard to come up with jokes, uh, with cybersecurity. It's such a serious field otherwise.
Evan Peña: No.
Ashish Rajan: To find moments of some kind of humor, it's almost like it's the driest humor you find.
Evan Peña: [00:41:00] Yeah.
Ashish Rajan: So I, I appreciate you guys smiling for my jokes. Of course. I appreciate you guys doing that. Where can people know more about the work you guys are doing and connect with you to basically talk more about automated pen testing?
Ashish Rajan: What's the website for Armadin? What's the LinkedIn?
Evan Peña: Armadin.com.
Ashish Rajan: Armadin.com. Yeah.
Evan Peña: Or you can go to hyperattack.ai
Ashish Rajan: and see. Oh. Yeah. You got that URL? Yeah. Hyperattack.ai?
Evan Peña: Yeah.
Ashish Rajan: That is a steal. I imagine you're gonna get a lot of people going, "Hey, I'll pay you money." Uh, that's a great URL.
Ashish Rajan: So hyperattack.ai and armadin.com. Armadin.com. Armadin.com. And are you guys on LinkedIn as well? I can put the links in. We're on LinkedIn, yeah. I, I'll put the links for LinkedIn as well. But thank you so much for coming on the show.
Travis Lanham: Thanks for
Ashish Rajan: having us. Sharing all that as well.
Travis Lanham: Thank you. This was wonderful.
Ashish Rajan: Yeah, thank you so much. Appreciate
Travis Lanham: it.
Ashish Rajan: Thank you. Thanks everyone tuning in as well. Thank you for watching or listening to that episode of AI Security Podcast. This was brought to you by TechRiot.io. If you wanna hear or watch more episodes of AI Security, check that out on aisecuritypodcast.com. And in case you are interested in learning more about cloud security, you should check out our sister podcast called Cloud Security Podcast, which is available on cloudsecuritypodcast.tv.
Ashish Rajan: Thank you for tuning in, and I'll see you in the next episode. Peace.

.png)
.png)

.jpeg)

.png)













