
Frontier models and the future of cyber defense.
Maria Varmazis: It'll come as little surprise that the two letters on everyone's lips at this year's Black Hat USA 2026 were, indeed, "AI." There's been a lot of focus lately on how AI is and can be used by attackers to exploit vulnerabilities and deploy social engineering campaigns at blistering speed and scale. But today, let's take a moment and turn our focus to the defenders. Welcome to this CyberWire special edition. Today we are sharing a conversation recorded at Black Hat 2026 in Las Vegas, where Dave Bittner sat down with Clint Gibler, Cyber Lead at OpenAI, and Robbie Winchester, Chief Global Professional Services Officer at SpecterOps, where they explored how frontier AI models are changing the way defenders approach cybersecurity. Their conversation moves beyond the hype to examine responsible AI deployment, AI red teaming, reducing noise in security workflows, and the balance between advanced models and human expertise. They also discuss OpenAI's Trusted Access for Cyber program and what it takes to give security practitioners access to powerful AI capabilities while managing the risks of misuse. Here's their conversation. [ Music ]
Dave Bittner: Hello, everyone. Welcome, and thank you for joining us here today at Black Hat 2026. We're delighted that you took some time out of your busy schedule to be with us here today. My name is Dave Bittner. I am the Producer and Host of the "CyberWire" podcast, and it is my pleasure to be your moderator here today with our guests for our panel discussion today. I have Clint Gibler, who is the Cyber Lead at OpenAI, and also Robbie Winchester, who is the Chief Service Officer at SpecterOps, who, of course, are our hosts here today. To kick things off, before we dig into the real meat of our conversation here today, can we just, kind of, do a level set when it comes to where we think we stand when it comes to our journey with tools like ChatGPT and OpenAI and the large language models? Where do you think we are in that journey?
Clint Gibler: I think a few years ago, or say, ChatGPT 3.5, it was an interesting idea, and there were, sort of, the glimmers of something practically useful and valuable there, but I think today, across many domains, whether it's coding, knowledge work, or other things, it is, at least for me, actively very, very helpful, surprisingly so --
Dave Bittner: Yeah.
Clint Gibler: -- and we see it continuing to improve. How about you? What do you think?
Robby Winchester: Yeah, I think we're, kind of, at an interesting place where there's this breadth of capability and advancement within the models and what LLMs can do, and the accuracy, and the kind of integrations that exist. It's also, at the same time -- kind of tangentially related to your question, there's a, kind of, learning how to adapt and adopt, and what is the right way to use this? There's, I think, a lot of, at the base, "I'm going to use this as a magic answer machine." You know, the, kind of, stereotypical, you ask questions, you get a response. But that's also only the really surface-level utility of, how can you integrate and use different capabilities and make, you know, kind of full-featured projects, things, and go beyond, like, replacing Google Search? I'm interested and excited to, kind of, see that shift. I think we're living that right now, where everyone knows I can ask ChatGPT instead of Google and get a question response, but not everyone knows how can I go and leverage something like Codex and do a project? You know, what are these different, like, next-generation, kind of, things, and what does that future look like, which is exciting?
Dave Bittner: Yeah, just quick for the audience, let me see a show of hands. How many of you are using these kinds of tools on a daily basis, would you say? Pretty much everybody, at least once a week? Yeah, everybody else. Now keep your hands up. How about a year ago? Would you say you were using them daily a year ago? Maybe half, so we can see how these have become part of our everyday use for so many people, and obviously, this is a room full of folks who are a little biased towards that kind of use, but it is interesting to see how people are coming to depend on it in their everyday lives, people you wouldn't expect to be calling on that. I want to ask you about the Trusted Access for Cyber program, which is giving vetted people access to these frontier models. Describe that for us and why that's an important part of how you're presenting these tools.
Clint Gibler: With regards to many of the things we do at OpenAI, it's useful to understand it from, sort of, a key point of view or a key frame of reference, and that is how do we maximally support and augment defenders while giving, ideally, minimal capabilities for attackers? So how do we protect the world rather than giving, potentially, very capable cyber models to attackers? Specifically, we have our mainline models, such as 5.6 Sol, that are very capable at a number of defensive tasks, whether that's looking for bugs in source code to analyzing logs to taking a look at malware -- as was shown earlier today -- so many different things. In the capability spectrum, from, you know, identifying and fixing vulnerabilities all the way to, say, generating exploits, which has a higher attacker uplift, the earlier things on that spectrum, we want as broadly available to as many defenders as possible. There are some trusted companies, or some trusted individuals, that do get a lot of value from a broader range, including more offensive capabilities, such as for AI red teaming or pen testing, which does have a defender uplift, but -- sort of, commensurately, a higher attacker uplift. So basically, the Trusted Access Program is our ability to give the most capable cyber models to defenders so that they can help secure both companies and their customers.
Dave Bittner: How do you calibrate that, and is there any collaboration among the leading providers of these sorts of tools with all the frontier models? Is there -- collusion is the wrong word, but is there a -
Clint Gibler: Collaboration.
Dave Bittner: Collaboration, thank you. Thank you.
Clint Gibler: Right on the tip of your tongue.
Dave Bittner: Yeah, I did. I did. I guess, you know, there's a certain sense that I think some people have that those frontier models, the experimental ones, we might be playing with fire a little bit, so we need to be careful that we don't get burned. How do you calibrate what's ready for the general public versus what we need to keep an eye on to make sure that it's ready?
Clint Gibler: It's a good question. I think there's a lot of new developments with frontier model capabilities that, you know, we're all figuring this out for the first time together. There is the Frontier Model Foundation, where all of the foundation labs, as well as other government organizations and, say, NVIDIA and other companies where we're getting together to determine, you know, what should our policies be more broadly, so it's not just a single company making decisions in isolation? All of us -- speaking for at least one foundation lab -- we think very carefully about what we can do to empower and augment defenders, ideally maximally augmenting defenders, while giving minimal attacker uplift. There's many different, sort of, nuances and small choices that you can make there, but I think all of the labs, to my knowledge-- I have friends at all of them. I think people are very earnestly doing what they can to help secure the world as quickly as possible, given the increasingly capable open source models.
Robby Winchester: I do also think, to kind of piggyback on that, part of the part of the challenge of this is if you look at-- and kind of from the SpecterOps, coming from a more of an offensive security, like, adversary perspective, how can this be used intentionally in a way to demonstrate what bad looks like so that it doesn't happen, you know, inadvertently? If you look at, like, the history of red teaming, penetration testing, anything, you are doing things that would-- you go and do a penetration test. You are performing illegal actions against a company that is only okay because they said it's okay, but under any other circumstances, a bunch of -- you know, my group of testers, if they did it one week later, they would be committing a crime. It's because we've agreed to go and do this adversarial thing under these specific circumstances that there's merit. The reason why you do that is because there's concept and there's reality. You can have a perspective of, what is this going to look like? Once the rubber meets the road, you actually can then, kind of, unpack that and see. I think a lot of the challenge is figuring out what are those, kind of, intentional, unintentional? You know, dealing with fire, like, fire has many use cases that you may not think about in the first place, and some of those may be more dangerous than you realized. You can't conceptualize every bit of that, so having that, I think-- like, from our perspective, having the ability to work with a, you know, less-restrictive model, to be able to identify maybe use cases, issues, concerns, considerations in an authorized way for the benefit of the company and the providers is much better than you find out about it the first time online.
Clint Gibler: Yeah, and like with many things, fire being a good example, sort of, fire can both keep you warm and cook your food and also burn down your house. If you are an enterprise that has many, many potential-- say you have different scanners and it's, like, you have 10,000, maybe, vulnerabilities, if you can dynamically prove some subset of those are actually exploitable, that helps you prioritize where the real risk is. So again, our cyber-capable models can work with SpecterOps and others to, yeah, just try to fix the most important things faster.
Dave Bittner: Well, and for the security leaders here with us today, what is your advice for the responsible deployment of these tools within their own enterprises? From what you've learned, any tips/tricks/words of wisdom?
Clint Gibler: Sure. So there's a number of security controls that you can use with Codex, for example. So rather than running it in a full-access mode that allows it to just do anything without prompting you, there's also an auto mode, which, essentially, has another model running to evaluate the tool calls and actions that it is taking, so that it isn't doing something potentially dangerous. You can also provide a custom policy for that. You're, like, hey, in our environment, here's the set of things that we believe are safe and trustworthy. Also, depending on the use case, you can use the right model for the job. For example, if you are scanning your network or trying to identify potential vulnerabilities, you can use, perhaps, a mainline model that will refuse to actually exploit it. However, if you are trying to actually prove something is exploitable, then, perhaps, you could use a more cyber-focused model, and there is a number of additional security controls that we are working on and rolling out soon.
Robby Winchester: I would take, kind of, a more -- maybe more abstract, taking a step back. The biggest thing from, kind of, like, my perspective is there are elements, like what Clint talked about, that are very much, kind of, related to the specifics of an LLM deployment that you would want to, like, make sure is there. Also, there are elements that are no different than any other application or no other piece of software being deployed, and yes, there's a lot of capability. But, like, why are you deploying it? What access does it need? Who needs to have that access? What are all the features? There's a lot of elements where you can inadvertently -- if you inadvertently are provisioning too much access; you're having over-permissions; you're having it too broadly deployed; there's nothing inherent about the LLM or the model or the capability there that is wrong necessarily, but you're creating fertile ground for something that you probably don't want to have happen, happen. So it's, I think, a little bit of, like, think about, in essence, this, there's -- at a base level, this is another application that is going to be deployed for a purpose; albeit a more broad and, kind of, interesting, nuanced purpose-- but it is an application you're deploying for a purpose. There should be an understanding of what you're deploying; how it is deployed; where it is managed? What are the identities there? What potential new and interesting attack paths? You know, if you're -- deploying an agentic tool to your SOC is potentially great, but that also now, potentially, opens up an opportunity for something else to take over your environment. That doesn't mean that you shouldn't do it, but you should have that perspective of, holistically, like, what am I adding and what is this, kind of, new landscape, rather than just adding a feature in? Again, any application has cost. Everything has risk. Making sure you're conceptualizing more of a complete view and not just focusing on, I think this can make solving a problem easier, and not, maybe, what's the second- or third-order effect?
Dave Bittner: Yeah, go ahead.
Clint Gibler: Oh, and I was going to say, we were actually talking right before this about how many core fundamental security best practices are still true and perhaps even more true. For example, if you are worried about an agent taking some action on either, say, a developer's laptop or someone in finance's laptop. You're like, oh, well, they could do this, and that's dangerous. Well, fundamentally, an agent running on someone's machine can, perhaps, have access to their credentials, their sessions, and things like that. So really, just, fundamentally limiting the capability of a single entity or, like, person or identity for taking some potentially dangerous actions. Like, that's something you should be doing anyway, and it's not necessarily different that it could be a human or an agent.
Robby Winchester: Yeah, the snowball can-- it's almost like you can more easily-- you always can start an avalanche, but this is potentially making it easier for that snowball to start an avalanche, so just being cognizant of where that may be the case.
Dave Bittner: Do you guys understand, like, it seems to me, like, there are -- obviously, there is great enthusiasm for these tools. Just walk around the show floor, and it's evident, but at the same time, there is a wariness, and I think in my mind, it seems justified. You know, we see news stories about the models breaking out of their sandboxes. So if I'm an enterprise security person, I'm thinking, am I -- you know, "Is this the Frankenstein monster I'm building in my environment, and how do I make sure I keep control over it?" because there are potential unknowns.
Clint Gibler: Kind of to feed off of that, part of the counterpoint that's in that same vein, but is a, it's going to happen anyways. It's already happening. There is going to be some unknowns for sure, but those are going to exist regardless, and it's kind of that element of, you know, you're already connected to the internet. You already are having different things. Every organization is there. I would say, probably, every organization is touched by AI and LLMs in some fashion, either knowingly or not, either directly or not, and so that dependence, and that, like, kind of, "system of systems." We have this big, interconnected society. We have this big, interconnected, like, environment. Other people or other organizations or risks existing, like, in that space are going to trickle down to some fashion. So my counterpoint to that is almost, you're going to have to go down that road anyways, and so why be avoidant? Try and at least confront it. It's going to always be a little bit new and interesting and scary, and it is going to still have a lot of unknowns, but you can't keep yourself out of the fight for a while. You're kind of already in it.
Robby Winchester: Yeah, along those lines, one potential analogy is back in the day when the cloud was very new, there were many security professionals and IT organizations more broadly who thought, you know, the cloud is scary. We can't control our machines. You know, clearly, we could never move off-prem. Then eventually, today, most companies -- or at least many companies -- do predominantly use AWS, GCP, and Azure. I think that in the early days of the cloud, it was not clear what security controls do we need? What security primitives do we need? Like, how does one operate in the cloud securely because there just wasn't precedent? I think, similarly, we're in a similar space with AI adoption. I think we're all figuring it out together, and I think we are speedrunning a lot of the best practices, security controls, and other things that just need to be invented along the lines of governance, observability, monitorability -- tons of work by great startups in this space, as well as labs and bigger companies. But I think, to your point, just the productivity gains and other benefits will probably outweigh security concerns.
Dave Bittner: Isn't that a fundamental difference? What you say is the velocity, right? We have this phrase now, "machine speed." "We need to defend at machine speed," and is that a new horizon? Is that different from the cloud adoption, just how quickly -- and it seems to be accelerating?
Robby Winchester: Thinking about it in machine speed is interesting because it also, kind of, frames it in a, there's a lot of perspective that this is creating an entirely new set of problems, and I think the maybe not common takeaway I have from machine speed is it's highlighting, there's a lot of things that you should have already been worried about that you just need to worry about more or faster. So, like, some of the basics of -- and I think it's highlighting, like, if you're detecting and responding at machine speed, part of that is you're trying to, like, almost fight fire with fire. Obviously, kind of, from a SpecterOps perspective, like, we're very big on the, kind of, attack graph, attack path perspective. What are the things? Like, how can we go and see what bad opportunities exist and try and take more of a preventative mindset? Fortunately, I guess, like, we kind of had that perspective prior to the proliferation of the LLMs, and that seems to, kind of, be a durable perspective today of can we mitigate? If we're talking about the common term or the phrasing now as, like, you know, initial access vulnerabilities are cheap. It's easy to get that initial access, because you can just hit everything all the time, phishing emails, like, all kinds of stuff. We're going to find vulnerabilities. That doesn't necessarily mean every step of everything is going to also be easy, so are there preventative mindsets? Can you take that machine speed? Like, can we also mitigate so that there's less things to have occur at machine speed, and not just think about it of, like, AI fighting AI. Then is there a peeker's advantage, like, Counter-Strike style of you want to hit the lag spike so you can get the shot off first?
Dave Bittner: Right.
Clint Gibler: There are, perhaps, some differences, but I think there is, also, a lot of, still, similarities, things that have always been true. You know, companies have had a number of vulnerabilities in third-party dependencies, as well as their first-party code, so that's always been true. It's still true. The ability for, perhaps, less skilled actors to find more of them faster, like, that's increased a lot. So the, like, "Eh, maybe we'll patch in a few months," or "Eh, this probably won't affect us," I think that assumption has changed. There's also -- I think I saw a blog post recently by someone who was just testing, like, okay, could an agent go into a cloud environment and pivot across several things and steal PII, and what's the efficacy? What's the speed? Things like that, and I think one interesting takeaway from his work to me was that, you know, I think CloudWatch and other logs from cloud providers, if those are only coming every five minutes or so, but you could do meaningful attacker actions, and under the, like, even logging window, I think that is, maybe, a meaningful difference. But broadly, being able to patch software quickly, being able to roll out changes to production and validate, you know, will this break things or not? Like, all of those core engineering best practice capabilities, I think, are still important and perhaps more important now, similar to using multi-factor. Like, all these things are just maybe more true now, but they're not meaningfully different, necessarily.
Dave Bittner: It doesn't change things; it unmasks them.
Clint Gibler: Well, and I think to that line, it's always been that question of, like, "Are you lucky or are you good?" Because you're never -- unless you are paying for a penetration test or you're having, like, regular things, like, if you don't get -- you know, if your company is not getting hacked, is it -- were you lucky, or were you actually stopping all of that? I think that's-- with the advent and capability and everything, it's potentially shifting that to where it's more important to be good, where maybe you could have gotten away with a lot more luck in the past.
Dave Bittner: Interesting.
Robby Winchester: Yeah, and along those lines, in terms of, say, cyber criminals, if you are trying to commit some ransomware or something like that, you have, you know, a fixed amount of human time. You, perhaps, have a certain number of people who can take those actions. So certain organizations may have not been worth your time in the past, but if the amount of time you need to just continuously probe and look for holes, like, if that cost and effort goes way down, then the set of targets that actually make sense to look at goes up. So yeah, in terms of, "Are you lucky or are you good?" I think companies that may have not been worthwhile to be targeted in the past, just because of the cost and economics have changed, might get more attention now.
Dave Bittner: One thing I want to make sure that we take some time for is red teaming. You mentioned red teaming earlier. Can we really dig into that? I mean, what are the elements of red teaming that the LLMs lend themselves to? I'm just curious of both of your thoughts of applying these tools to that specific task. Maybe I'll start with you, Robby, because I know that's sort of your jam, right?
Robby Winchester: Yeah, so I think there's two facets of it that are both very interesting. One is what potential capability and support can be there? Like, what tools? What is the -- you know, in the same way it is easier for you to write an email that sounds all professional for work, it's also easier potentially for a -- you know, spelling errors and grammatical problems should be a thing of the past. So there's a side of it adding capability, being able to, like, find exploits, develop potential malware, rapidly iterate on capability, you know, help process, like, large amounts of data to kind of go through the attacking problem set. Again, the, like, having trusted cyber -- like, not having those refusals makes that more of a, kind of, stream, and as you have the potential proliferation of, like, open models, that may become the case. Like, if you don't control that, there may be no guardrails, so that's, like, being able to see what that perspective is to provide it. The other side is the, what attack surface does it open up? There's AI red teaming in a, how can we use it to red team? But there's also, do you know what you're implementing? Like, do you have a -- are you getting your models from a controlled place? If you're rolling your own models, are you doing types of evaluations? If you're doing, you know, RAGs or other types of internal, like, supplementation, like, is that a controlled process, or can anyone just go and touch that? You know, can I take over your enterprise? Like, I can take over your enterprise GPT account, for example, and now I have all this access where, without that, you wouldn't conceptualize that initially. So I think there's a little bit of that reframing of there's the ability, and then there's the potential, like, every attack surface. Like, what attack surface are you adding and having, kind of, a joint fusion of that perspective?
Dave Bittner: Anything to add, Clint?
Clint Gibler: Yeah, a couple of things. Building off of what you were discussing regarding capabilities: I think, historically, security professionals have been great at having deep security SME expertise, but maybe we're not the best developers. But now, both on the positive side, like, application security or product security engineers can now build, like, secure-by-default libraries and tools and infrastructure that really scale security programs. The other side of that is red teams or threat actors. Now, what they may not have had the capacity to write very detailed or interesting offensive capabilities, now they can because they're very capable coding models, or they could build, you know, custom things for a target, so I think that capability is higher. Another thing I would say, maybe just one concrete, interesting use case that our internal red team at OpenAI has played around with is in your environment, you have a series of security controls that you expect to be working and providing certain protections, whether that's sandboxing or perhaps network isolation. You know, there's many classes of these. So having, sort of, continuous red team agents placed at different places within your environment which are specifically tasked with bypassing or evading those controls. Like, you can imagine having invariants about your environment which you assume are always true. For example, in this subnetwork, you should never be able to talk to this database or, you know, you should never be able to reach the internet from this place. Like, there's a bunch of these properties or invariants that you might expect to hold when you're designing the environment and building it. Sometimes those aren't true, and models are very persistent and capable. I think one thing that we have already been doing and are continuing to do is just continuously testing all of these security properties. I could see other companies doing something similar.
Dave Bittner: Right, just tirelessly banging away, checking, and testing and making sure over and over again.
Clint Gibler: Yeah, like, we assume we have many layers for this, but let's make sure that's true.
Dave Bittner: Yeah, yeah, all right.
Robby Winchester: Yeah, consulting can be frustrating.
Dave Bittner: [laughter] All right. Well, Clint Gibler is Cyber Lead with OpenAI and Robby Winchester is Chief Services Officer at SpecterOps. Gentlemen, thank you so much, and thanks to all of you for joining us here today. I appreciate it.
Clint Gibler: Yeah, thanks for having us. [ Music ]
Maria Varmazis: Our thanks to Clint Gibler, Cyber Lead at OpenAI, and Robby Winchester, Chief Global Professional Services Officer at SpecterOps, for sitting down in conversation with Dave Bittner. I'm Maria Varmazis. Thanks for joining us today. [ Music ]

