Search
 

OT Incident Response: Critical Lessons from the Field

Confirmed OT incidents are still rare compared to IT, but that number says less about risk than most people assume. In this episode, recorded while Matt Calligan was in Des Moines speaking at a fusion center event, Matt sits down with Chris Sistrunk, Technical Leader on Mandiant’s ICS/OT Security Consulting team and a former senior engineer who oversaw transmission and distribution SCADA at Entergy for more than a decade. Chris has responded to many OT outages, helped stand up the ICS Village at DEF CON, and is one of the field’s most respected voices on a threat that is only half a joke: squirrels. The conversation covers why OT incident counts stay low and what is quietly changing, the all-hazards reality where weather, equipment failure, and human error still outnumber nation-states, what digital forensics and incident response actually look like when safety outranks evidence, how responders keep teams talking when the primary network cannot be trusted, and why cyber-informed engineering treats security as a design discipline rather than a bolt-on.

Listen on :

  1. Low OT incident counts don’t mean low exposure. Many incidents go unrecognized: human error that let malware in, PLCs accidentally exposed to the internet, and commodity malware sitting undetected for thousands of days until something breaks. Nothing bad has happened, so no one has looked.
  2. Connectivity is eroding OT’s old separation. More cloud, more AI, and “a brain and a network connection in just about everything” mean more IT attacks bleeding into OT, as with Colonial Pipeline. It is getting hard to find a purely analog device.
  3. Squirrels are the real-world baseline. In Chris’s view, there has not been a North American power outage caused by a cyberattack; weather, animals, equipment failure, and human error drive most outages. When the lights went out at the Super Bowl, everyone asked “was it cyber?” The cause was a relay design defect.
  4. AI empowers the middle, but intent still limits the threat. Tools like AI let lower-skilled actors build things (a DNP3 scanner, PLC code) that once required nation-state or well-funded ransomware resources. Chris notes intent still matters: attackers rely on the same power and water, which confines most cyber-physical attacks to advanced, remote adversaries, though lone wolves remain possible.
  5. The basics still beat everything from script kiddie to APT. Get things off the internet, minimize exposure, use two-factor everywhere, keep backups and critical spares, and have an OT incident response plan. Chris’s line: publish a 20-year-old article today, change only the date, and it would still apply.
  6. The first hour of OT response is human. Chris often acts as a counselor before a forensic analyst: nobody died, the water is flowing, OT is segmented, now get some sleep and some food. Teams calm down when they treat a cyber incident like the storm restoration or incident command they already know.
  7. In OT forensics, safety gets a vote. Many IT forensics techniques still apply, but you cannot pause a controller managing molten steel to take an image. Plant managers, engineers, and safety staff weigh in, and sometimes the right call is to swap equipment and return to a known state instead of preserving evidence.
  8. Don’t reinvent incident communications. Borrow the war room, stand-up cadence, and ICS/NIMS structure the business already uses for hurricanes and ice storms, and pre-agree back channels off the compromised network. Chris’s old utility kept radios as the backup to cell phones, and satellite phones as the backup to the backup.
  9. Cyber-informed engineering bakes security into design. CIE works backward from consequences and failure modes to design resilience in, rather than bolting it on. Physics sets real limits (encryption can’t slow a 100-millisecond trip command), so a secure device is one component of the whole system, not the whole answer.

Navroop Mitter:

[00.00.03.21–00.00.29.18]

This is Navroop Mitter, founder of ArmorText. I’m delighted to welcome you to this episode of the Lock and Key Lounge, where we bring you the smartest minds from legal, government, tech and critical infrastructure to talk about groundbreaking ideas that you can apply now to strengthen your cybersecurity program and collectively keep us all safer. You can find all of our podcasts on our site and listen to them on your favorite streaming channels.

Navroop:

[00.00.29.20–00.00.33.06]

Be sure to give us feedback.

Matt Calligan:

[00.00.33.08–00.01.00.10]

Welcome again to the show. I’m Matt Calligan with ArmorText, coming to you from the modern Midwest here in Des Moines, Iowa. I’m actually bringing you this episode while speaking at a fusion center event. And in this episode, we’re getting back into OT territory. But from the perspective of a practitioner who has responded to many OT outages, My guest in this episode is the most understated intellectual powerhouse you’ll meet in the digital forensics and incident response field.

Matt:

[00.01.00.15–00.01.24.03]

And we’re going to cover things like what digital forensics and IR actually look like in OT, why OT incident rates are low and what’s changing, and the all-hazards approach, which includes malicious state actors, sure, but also squirrels and human error, as well as the current state of nation-state activity in OT. My guest is Chris Sistrunk.

Matt:

[00.01.24.04–00.01.48.06]

Chris is the Technical Leader on the Mandiant ICS/OT Security Consulting team, which is part of Google Cloud. He’s been there for the better part of a decade. Before Mandiant, he was the Senior Engineer at Entergy for well over a decade, overseeing the transmission and distribution SCADA systems. Chris helped stand up the ICS Village at DEF CON and co-founded the BEER-ISAC, which is my favorite

Matt:

[00.01.48.07–00.02.00.12]

ISAC. He’s also one of the most respected voices on the topic of the dangerous world of squirrels. So with no further delay, let’s get into the discussion. Chris, welcome to the show.

Chris Sistrunk:

[00.02.00.14–00.02.02.00]

Thanks for having me, Matt.

Matt:

[00.02.02.05–00.02.04.05]

It’s good to see you again.

Chris:

[00.02.04.07–00.02.05.20]

Yeah.

Matt:

[00.02.05.21–00.02.40.05]

All right. Let’s dive right into the first segment, then, about the incident ratio paradox that still exists out there. I have some thoughts on it, but I’m sure you do too, and that’s what I wanted to get into. So, you have been doing OT security for decades. One thing that tends to surprise people who have been on the IT side, or aren’t aware of either side, is how few confirmed OT incidents there are compared to IT as a ratio.

Matt:

[00.02.40.05–00.02.56.08]

And in your experience, why has that number been so low so far? Do you see that changing, and what lessons can we learn, if there is a transition or not, from your perspective?

Chris:

[00.02.56.14–00.03.20.03]

Right. And I don’t know if there’s a specific exact percentage; it’s not known, the ratio that you’re talking about. But we have what we’ve seen, what others have reported on, and OT-specific incidents are not as common as, like, IT, cloud and others, right.

Matt:

[00.03.20.05–00.03.21.10]

Dramatically, though. There’s like…

Chris:

[00.03.21.11–00.03.22.22]

Yeah, yeah.

Matt:

[00.03.23.00–00.03.24.15]

…an order of magnitude difference.

Chris:

[00.03.24.16–00.03.58.09]

Yeah. But there’s a trend, kind of in the last several years, where there have been a lot of IT attacks that have bled over into OT, that have had direct or indirect impacts. Like ransomware, anything targeting, say, Windows or a firewall brand or whatever. You’ll have those impacts there. So that’s the kind of thing that we’re seeing, right?

Matt:

[00.03.58.11–00.04.12.15]

I mean, Colonial Pipeline has been public about this, but it was an IT attack, and the shutdown impacted the OT, right? As far as I know.

Chris:

[00.04.12.20–00.04.31.17]

That’s right. And there’s a lot of context. You can go listen to all of the discussions around that. Some of it’s on C-SPAN, where they had hearings and things like that. You can go get the details. But yeah, absolutely.

Matt:

[00.04.31.18–00.04.48.20]

Do you see that changing? Do you see evidence that the ratio is going to narrow? Like 100 water utilities getting hit, the PLCs getting hit, clearly low-hanging fruit. But do you see reason to believe that shift is happening?

Chris:

[00.04.49.02–00.05.15.14]

I would say that there’s a shift in technology, where companies are using more cloud and definitely more AI than ever before. There are a lot of businesses that want to use it and have found value in using it, or they may be forced to use cloud. I know there’s a lot of technologies that are no longer available on premise, right?

Chris:

[00.05.15.15–00.05.42.04]

And even Windows requires an internet connection for the latest versions, things like that. Now, there are ways to make Windows 11 work offline, but I’m just generalizing here. So for anyone listening, yeah, I know I’ll get arguments against me. But those are some of the themes that we’re seeing. We’re seeing more people than ever connect things more than ever.

Chris:

[00.05.42.06–00.06.07.19]

As a large blanket statement. Now, yes, there is still going to be OT that is on prem only, or a hybrid approach where it has on prem plus cloud or internet connections. But we’re seeing more connectivity than ever before. You go to every trade show, it’s there. The AI is there. Every kind of trade show, it’s there.

Chris:

[00.06.07.19–00.06.28.09]

The new technology is shifting. So it’s like the dawn of the internet, right? And then the people who didn’t adopt the internet didn’t shift. It’s like Best Buy versus Netflix, right? One adopted the internet and one didn’t.

Matt:

[00.06.28.09–00.06.30.23]

Or Blockbuster, right?

Chris:

[00.06.31.04–00.06.36.00]

That’s what I meant. Not Best Buy; I meant to say Blockbuster.

Matt:

[00.06.36.01–00.06.55.16]

Exactly. The innovator’s dilemma kind of argument. So from your seat, you’re seeing connectivity requirements for the latest technologies kind of breaking that barrier, or jumping that barrier, creating new access vectors.

Chris:

[00.06.55.18–00.07.12.05]

Right. And now there’s a brain in everything, from sensors all the way up to the latest IoT devices. There’s a brain and a network connection, or a wireless connection, in just about everything.

Matt:

[00.07.12.06–00.07.12.17]

Yeah.

Chris:

[00.07.12.18–00.07.30.14]

And it’s almost impossible not to find something like that. It’s almost impossible to find pure analog devices that are not connected right now. Again, that’s a generalization, so those of you throwing tomatoes at me, I already know that.

Matt:

[00.07.30.16–00.07.56.12]

This is a safe space here, Chris. You can get spicy and punchy all you want; that’s what we’re here for. So, I was reading the other day, when CERT Polska put out their analysis, and one of the things that stood out to me was that the initial alert came from an engineer, but the engineer classified it as human error.

Matt:

[00.07.56.12–00.08.19.12]

And it was actually CERT Polska who said, wait, that looks a lot like something different that we’re seeing. They chose to investigate it anyway. Have you personally seen evidence of underreporting of incidents, because people don’t know how to recognize them as something other than a benign issue?

Chris:

[00.08.19.14–00.08.50.19]

Absolutely. We see instances where an honest mistake, human error, has let malware into an environment, or just exposed something to the internet. I’ve talked to water utilities who had a setting wrong, and it exposed their PLCs to the internet. Like there’s a there’s a lot of these different things. There’s also folks that didn’t know they had an incident.

Chris:

[00.08.50.20–00.09.20.00]

And the dwell time is in the thousands of days. It’s like it was 15 years ago for regular APTs. They just didn’t know they were in there. And sometimes a control system owner will have malware, like commodity Windows malware, and not know it until the network is saturated and they have to find out what’s wrong.

Chris:

[00.09.20.00–00.09.43.17]

So there are a lot of underreported incidents like that. Do they have impacts? Probably not, because nothing bad has happened. So it’s just living there. Not to say that something bad couldn’t happen. It just hasn’t happened, so no one’s looked, right?

Matt:

[00.09.43.18–00.09.48.23]

Right. It’s not being recognized, just because nothing’s broken yet.

Chris:

[00.09.49.05–00.10.04.22]

And people go to the doctor. They only go when they know something is wrong. And in some cases, like old farmers: I’ll wait till my arm’s cut off before I go to the doctor.

Matt:

[00.10.04.22–00.10.26.12]

I can’t swallow. Well, you probably should have looked at that a long time ago. As far as incidents themselves, you’ve talked about squirrels a lot, right, maybe being a bigger concern than nation-states. And it’s always good for a laugh, but unpack that part

Matt:

[00.10.26.13–00.10.32.22]

as far as the non-malicious piece. It’s a joke, but it’s a real component of it.

Chris:

[00.10.32.23–00.10.54.14]

It’s a joke. You can’t see this on the podcast, but I’ve got a patch: CyberSquirrel1, our number one nemesis in the power grid. For those in the power grid, I mean, those are reality. There have not been any electric power outages in North America due to a cyberattack. That just hasn’t happened.

Chris:

[00.10.54.15–00.11.29.09]

There have been attacks and there has been malware, but there haven’t been any power outages. And so that kind of puts things at a level set. We have more problems from bad weather, Mother Nature, and that includes squirrels, equipment failure and human error. Those are the main ones. Like, when I was at the power company that had the Super Bowl power outage, they were like, oh, was it cyber?

Chris:

[00.11.29.10–00.11.49.01]

I’m like, I don’t know yet. And this is even a decade ago: oh, it had to have been cyber. People want to jump to that conclusion. And I go, no, it’s more likely a squirrel or a rat bit into something. In that case, it was a relay operation.

Chris:

[00.11.49.02–00.12.01.04]

And it was a defect in the design of that relay. And there were a few small misconfigurations there, but it wasn’t malicious in any shape or fashion.

Matt:

[00.12.01.05–00.12.01.22]

Right.

Chris:

[00.12.02.01–00.12.34.04]

And so that’s kind of the context around cyber incidents, not just for the power grid but for control systems in general. Failures happen all of the time. Water main breaks happen all the time, and most people don’t even hear about it. There are some places that don’t even know that a problem has occurred, because these things have been happening a long time, and engineers and operators know how to fix them.

Matt:

[00.12.34.05–00.12.52.07]

Right. I’m assuming, though, the implication here is not that we should. There’s designing a security model for the most likely event, and then there’s a different security model you build for the event that has the largest impact.

Chris:

[00.12.52.08–00.13.22.16]

Right. That’s the low-frequency, high-impact events, and then the high-frequency, low-impact events like this, that whole continuum there. But in this industry, cybersecurity in general and especially for grids, we’ve been practicing this for, gosh, 20 years now. And there’s a whole cadre of people working on this problem in every facet. And there’s a lot of hard work that has happened.

Chris:

[00.13.22.16–00.13.45.22]

When I first started in OT security, I guess I can blame Stuxnet. In 2010 I was still at the power company, and it got me excited and scared about cybersecurity. We’ve all come into this. That was a small community back then, and now it’s a much, much larger community. But we still have a ways to go.

Chris:

[00.13.45.23–00.14.05.04]

Obviously there are the cyber poor that are out there, like the water utilities that have 2,000 customers and five employees. They don’t have a system and they don’t have an OT security program. They’re just trying to make sure people have water and are able to flush their toilet.

Matt:

[00.14.05.06–00.14.06.09]

You know. Right, right.

Chris:

[00.14.06.10–00.14.16.08]

And so we’ve got a lot more work to do. But it’s good to look at things in context.

Matt:

[00.14.16.09–00.14.26.21]

Sure. Within that low-likelihood, high-impact scenario, usually nation-states are the most common. Is that…

Chris:

[00.14.26.23–00.14.48.20]

Yeah, that’s what we’ve seen. There have only been a few cyber-physical attacks that have had any physical effect, like Stuxnet, or the cyberattacks against Ukraine that caused power outages, or

Chris:

[00.14.48.22–00.14.55.14]

Triton. And there are a few other minor ones. But again, they didn’t leave lasting damage.

Matt:

[00.14.55.15–00.14.56.08]

Right.

Chris:

[00.14.56.09–00.15.06.13]

Other than Stuxnet, and even then it didn’t completely stop the nuclear program.

Matt:

[00.15.06.14–00.15.07.06]

Right, right.

Chris:

[00.15.07.07–00.15.35.01]

So that’s some observations there. Not to say we shouldn’t keep working on it, in case there is a high-impact event. There have definitely been high-money-impact events that have cost a lot of money. But as far as loss of life, mega-disaster-level loss of life or loss of things, that hasn’t occurred.

Chris:

[00.15.35.02–00.15.51.05]

Now, there have been cyberattacks on hospitals where people have been impacted or died because of it, indirectly or directly. But that’s why we do what we do. A lot of us in the industry feel that way.

Matt:

[00.15.51.08–00.16.15.22]

A lot of security models, even when it comes to cyberattacks, are built around the fact that the high volume of malicious activity is low-hanging fruit by low-skilled actors. Do you see nation-states operating at a different level of skill set, or do they tend to still focus on low-hanging fruit as well?

Chris:

[00.16.16.00–00.16.52.00]

I think they do everything. If IT hackers are lazy, then even malicious hackers are going to be lazy. If something’s easy for them to grab, they’re going to grab it. But I would say now the lower-skilled attackers are being empowered by tools, like AI. I cannot code very well. That’s not my background. But I can tell Gemini to,

Chris:

[00.16.52.02–00.17.21.01]

hey, write a DNP3 scanner, or write this analysis tool. And it does a pretty good job. And so if you have someone with a malicious mindset, and somebody with access to an AI that doesn’t have guardrails, that can be a force multiplier for evil. And that’s something that, really, before now we didn’t have. Only nation-states had that kind of…

Matt:

[00.17.21.02–00.17.24.01]

…ability to invest in someone who could build that level of skill.

Chris:

[00.17.24.02–00.17.35.05]

Or rich ransomware gangs or groups or things like that, where there’s a lot of time and money invested in these tools.

Matt:

[00.17.35.07–00.18.04.14]

So, to paraphrase: the ease of access to a technical tool that can act basically as a proxy for a skilled hacker, for low-skilled people. Because there are a lot of people kind of dunking on AI doomers, right?

Matt:

[00.18.04.15–00.18.30.00]

AI is going to make everybody so much smarter and we’re going to see this massive shift. And they’re saying, well, there’s no evidence of it. All these things have happened, but there’s no change in scale of impact. Do you see that changing? Do you see these low-hanging-fruit, opportunistic smash-and-grabbers having a bigger impact with these tools in the future?

Chris:

[00.18.30.02–00.19.06.14]

Yeah, I don’t know. I think I read, and I’m having trouble remembering where I read this recently, but they used Gemini, or not Gemini, they used an AI tool to develop this. Oh, it was in the CERT alert about someone using AI to write Siemens PLC code. And I was like, well, if they had known, these tools have been existing. You could go download an app in the Apple Store.

Matt:

[00.19.06.17–00.19.07.23]

The Play Store, right?

Chris:

[00.19.08.01–00.19.32.11]

Yeah, that talks to Siemens PLCs. There are existing tools. And so some people are just using AI and not using their brain to go find the existing tools that do that. A tool can be used for good or evil; it’s the intent behind it. So it’s the same with the internet.

Chris:

[00.19.32.11–00.19.38.00]

It’s the same with software. It’s the same with AI.

Matt:

[00.19.38.01–00.20.09.07]

I guess the argument isn’t a restriction-of-AI argument. It’s, should we recalibrate our thinking about what’s the most likely kind of adversary? Sure, it’ll be squirrels. So far the math has been, well, until a nation-state sets its targets on us, we just have to worry about squirrels. But now we’ve got this middle ground of traditionally smash-and-grab opportunists, guys who now theoretically, and we’re starting to see it, move.

Matt:

[00.20.09.08–00.20.31.01]

The Monterrey water thing. And the LLM that won the capture the flag for Dragos. I did a podcast with an OT pen tester at Con Edison, and he said during a SANS workshop they took three or four screenshots of PLC ladder logic he had never seen written in that code before.

Matt:

[00.20.31.02–00.20.50.18]

Didn’t know a thing about it. And the AI tool reconstructed the entire code, found the vulnerabilities, told them how to get into it. So it’s like, if you know just enough, and like you said, that baseline information is out there in help manuals and things.

Chris:

[00.20.50.20–00.20.51.10]

Sure, sure.

Matt:

[00.20.51.11–00.20.59.22]

I’m just wondering, do you see a recalibration around recognizing that shift in the threat actor?

Chris:

[00.21.00.00–00.21.09.07]

Yeah, it’s kind of early to tell from my point of view, because I’m limited in scope to mostly OT. And so…

Matt:

[00.21.09.09–00.21.10.16]

I guess I mean OT-specific.

Chris:

[00.21.10.17–00.21.39.01]

Yeah. It’s back to intent. A threat actor likes electricity and water too. So why would I take out the power grid in my own area? I’d be shooting myself in the foot. So that means, okay, it’s someone across the globe that wants to create harm on someone else remotely.

Chris:

[00.21.39.06–00.22.05.04]

Without being caught, because it’s in a different country. Or, hey, they can’t extradite me here, because we’re in a safe harbor, an evil country that harbors criminals, right? So there’s some of that there. And there are also the governments out there that would retaliate after a cyber-physical attack in multiple ways.

Chris:

[00.22.05.04–00.22.30.02]

And I’m not going to get into those things because that’s not my expertise. But you can read about the different governments out there and their positions on these things. And so that’s kind of a barrier to most everyday types, maybe script kiddies, I guess, and it limits it to the more advanced attackers.

Chris:

[00.22.30.07–00.22.39.06]

But that doesn’t mean you won’t see a script kiddie who is crazy and wants to go after critical infrastructure.

Matt:

[00.22.39.07–00.22.41.02]

Yeah, yeah.

Chris:

[00.22.41.04–00.22.53.19]

Because we always have lone wolves and others that just do things because they can. I don’t know if that answers your question.

Matt:

[00.22.53.21–00.22.58.04]

There is no answer to it right now, I mean, that’s…

Chris:

[00.22.58.06–00.23.45.08]

That’s kind of what I’m seeing. And I think for us as defenders of OT, and the asset owners out there, if we stick to the basics and stick to what has been tried and true: get things away from the internet, make sure your exposure is limited to the absolute minimum, two-factor all the things, backups, having critical spares, and having a plan, an incident response plan for OT. Just doing those things will help protect against anything between the script kiddie and even APT.

Chris:

[00.23.45.10–00.23.59.12]

It’s something that a lot of people are starting to learn about, but it’s not new. We’ve been talking about the basics for a long time. I saw Kelly Jackson Higgins at Black Hat. I don’t know if you know her.

Matt:

[00.23.59.14–00.24.00.19]

She’s she’s.

Chris:

[00.24.00.21–00.24.13.10]

Yeah, she’s been at Dark Reading for 20 years. And I was like, Kelly, I bet you could take some of your first articles and re…

Matt:

[00.24.13.13–00.24.13.23]

…still apply.

Chris:

[00.24.14.00–00.24.23.21]

Publish them. Publish them today as is. It’s a 2026 headline. Change nothing but the date and it would still apply.

Matt:

[00.24.23.23–00.24.26.07]

Yeah. The concepts just don’t change.

Chris:

[00.24.26.08–00.24.32.15]

They don’t change. But it’s not sexy, because, oh, it’s AI threat defense, and we’ve got to do this, and…

Matt:

[00.24.32.16–00.24.32.20]

So.

Chris:

[00.24.32.20–00.25.19.17]

Right. So for the OT folks, if we do the basics and we do them well, we can just do our jobs, and the stuff that we make work every day in gas or water, wastewater, electric power, manufacturing, that will still be able to work as normal. Now, where we have more complicated environments, where there’s cloud involved or multiple sites relying on enterprise resource planning software, okay, now we have to figure out the critical data flows and protect those, and have a backup plan for when the internet goes out or the cloud goes down or whatever.

Chris:

[00.25.19.23–00.25.47.06]

And we have to be more resilient. So a lot of us in OT have been saying, hey, let’s not forget about resiliency and safety. Those concepts have been strong throughout safety and now OT security. So we can’t forget those things. And I guess I’ll be the old man yelling about this until I’m 100 years old.

Chris:

[00.25.47.07–00.25.48.11]

Right. It matters.

Matt:

[00.25.48.12–00.26.06.17]

What about the utility in Maine? I think it’s a co-op or a muni, but they actually got hit by China, and it was their adherence to the basics, to my understanding, as it reads, that kept it from turning into a full-blown incident.

Chris:

[00.26.06.18–00.26.36.05]

I don’t know the details about that one, but it’s one of those things where, I can’t tell you how many times we’ve seen a firewall that is decent enough to protect IT to OT, and it’s bought them enough time. The IT is getting hit with ransomware or whatever, and it buys them enough time to sever the network and to run their incident response plan on the OT side.

Chris:

[00.26.36.05–00.26.58.00]

Even if they didn’t have one, they kind of know what to do ad hoc, right? They say, okay, we can run this, we can run this. The plant manager is like, okay, yeah, this is our way to run disaster recovery mode. And even though it can be a stressful thing, it saves the company, or operations.

Matt:

[00.26.58.01–00.26.59.20]

Yeah. It’s always stressful.

Chris:

[00.26.59.22–00.27.30.01]

I feel like sometimes when I’m in these situations I’m more of a counselor. People are losing their mind. I’m like, it’s okay, we had an incident. Let’s break down the impacts. Did anybody die? No. Is the power or water flowing? Yes. Okay. Is OT segmented? Yeah. And I’m like, okay. Then you can see the level of…

Chris:

[00.27.30.05–00.27.52.11]

the level of stress starts to melt. And I’m like, are you getting enough sleep at night? No, we’ve been working 20-hour days. I’m like, no, you have to. When I worked for the power company, we had to have an eight-hour safety layoff, because line crews working, restoring power, that takes a big toll on the mind and the body.

Chris:

[00.27.52.12–00.27.59.06]

And then you also have to, I’d ask the team on the OT side, have you eaten?

Matt:

[00.27.59.06–00.28.01.08]

Well, right.

Chris:

[00.28.01.10–00.28.24.04]

And, no, we haven’t. We’ve just been having ramen in the office. I’m like, no, get somebody to bring in pizza. And I said, you already know how to do this. It’s the same as storm restoration or incident command. Stand up what you already know. This just adds a new wrinkle to it. And treat it like an all-hazards response.

Chris:

[00.28.24.04–00.28.54.11]

And they go, oh yeah, we’ve done this before. Let’s treat it like an ice storm. And now we have a framework that people already know. They start to feel better and they go, okay, now we can tackle this. But when the event first happens, if it’s never happened before, people can really freak out. It’s a mentally crushing thing to have a cyberattack at your company, for most people.

Chris:

[00.28.54.11–00.29.02.11]

And so they’re like, oh golly, what do I do? And this is why we’re here. We’re here to help and get you back on your feet.

Matt:

[00.29.02.12–00.29.11.05]

Is that first hour or two of your involvement typically what you’re doing, focusing on the human side of the equation?

Chris:

[00.29.11.08–00.29.34.12]

I mean, that’s what I’ve always cared about, because if you don’t have a good team, then the response is going to fall apart, right? It’s like when someone has a car wreck and there are a lot of people standing around: you call 911, you do this, you do this, you go get some road cones.

Chris:

[00.29.34.13–00.29.59.16]

That’s kind of me as an incident commander. We have to take charge of the situation if no one has done so, and be a good Samaritan, trying to help the situation. And I think a lot of my peers, within my company and across the world who do this job, would feel the same way.

Chris:

[00.29.59.18–00.30.17.23]

Because OT security is a very altruistic thing. We want to make sure that people’s water and power and food and all this other stuff still works like it does today.

Matt:

[00.30.18.03–00.30.46.15]

When you’ve moved through that calming, getting everybody on the same page mentally, sort of recalibrated, and you’re diving in: some of your talks, I think it was DEF CON, discussed some of the challenges of traditional vulnerability scanning and things like that in an OT environment.

Matt:

[00.30.46.15–00.31.09.20]

So what kind of constraints? Because I was just in a LinkedIn debate recently about this sort of executive, high-level IT framework trying to be pushed down onto OT, and it echoed some of your presentations about scanning and things like that when it’s an OT-specific environment.

Matt:

[00.31.09.20–00.31.13.15]

So talk to me about some of those differences that you’re seeing there.

Chris:

[00.31.13.21–00.31.36.05]

Yeah, that was my “what’s the difference” talk, the one I had at Black Hat, and the DEF CON talk was NSM for ICS. So I kind of had some themes there those years, from 2015, 2016. There are some things that absolutely translate from IT into incident response and forensics in OT. You have Windows, you have networks.

Chris:

[00.31.36.05–00.32.03.21]

That’s not going to change. What is going to change is, hey, we can’t take time to do a forensic image on this controller, because it controls, oh, it’s a steel mill and it’s got molten steel. We can’t let that stuff freeze, right? Otherwise you’d have to throw away the whole thing.

Matt:

[00.32.03.23–00.32.04.08]

Yeah.

Chris:

[00.32.04.09–00.32.32.03]

Yeah, a big chunk of metal. So there are some situations where you have to make a critical decision: do we want to capture the forensics, or do we want to make sure safety’s the priority? And so that’s where the experts, the plant manager, the general manager for that site, or the chief engineer, or the site engineer, or the lead operator, their input is going to be important.

Chris:

[00.32.32.06–00.32.59.23]

Same with the safety folks. Safety gets a vote in this. And so if you can get forensics in OT, that’s great. If you can keep from rebooting a PLC, to grab the memory out of it, if that’s even possible, it’d be great to do that. But in a life-or-death situation, or a critical safety issue, you can’t wait on that.

Chris:

[00.32.59.23–00.33.27.09]

You’re going to swap a piece of equipment out. You’re going to reboot something. You’re going to try to get it to a known deterministic state. And so sometimes the traditional things have to go on the back burner. Also, some of those tools don’t work in OT, obviously, protocol-wise. Although now there are a lot of tools out there that support OT protocols.

Chris:

[00.33.27.09–00.33.54.17]

And there are so many sensors out there, free and enterprise level, where it’s a professional tool. There’s Dragos, Claroty, all of the ones that are out there; I can’t name them all. And then you have issues where, oh, well, we can’t install agents because the system is so old. We’d like to put our own agent in those places, and

Chris:

[00.33.54.18–00.34.28.07]

we can’t do that in those cases. We’re running Redline on a Windows XP box to get memory, or Volatility or something like that. But if it’s a Windows system, if it’s virtual-machine based, or Windows 10 or newer, or Windows Server 2008 or newer, traditional IT forensics and incident response will absolutely apply, just guided by the operational requirements.

Matt:

[00.34.28.08–00.34.29.10]

Right, right.

Chris:

[00.34.29.15–00.34.30.11]

Yeah.

Matt:

[00.34.30.13–00.35.04.09]

When you’re actually on location, a substation or whatever, in a suspected incident and it’s ongoing, where extraction or remediation or removal has not happened yet: obviously at ArmorText we’re always talking about communications. But what kinds of ways do you see teams communicating, externally and internally, across different divisions?

Matt:

[00.35.04.09–00.35.16.20]

Maybe you’re not all in the same room together. How do those usually get stood up? How do they keep in sync with each other, even if they’re not all standing in the same room during one of these scenarios?

Chris:

[00.35.16.22–00.35.47.04]

The company may have a war room process that they have for any incident. If it’s an ice storm or a hurricane or a fire, they usually have that for their business continuity. And so we latch on to those existing processes. And if they don’t have one, then we can guide those folks into having a regular stand-up call.

Chris:

[00.35.47.06–00.36.12.01]

Then there may be one that’s just an open channel all day for people to come into. Some people follow the ICS, the FEMA NIMS process, where they’re doing the planning cycle, and they’re doing the whole shift change on who’s the incident commander, who’s the deputy, who’s the planning section, who’s the operator. All of those things depend on the cadence.

Chris:

[00.36.12.04–00.36.43.09]

We try not to reinvent the wheel here. Getting people to talk is critical. Getting updates from different teams is critical. And that’s what incident command is all about. So if you don’t know about the Incident Command System, or FEMA’s National Incident Management System, talk to your local EMS or fire or police. They all have to do it. And you can go to ICS4ICS and see all of the OT-specific things.

Chris:

[00.36.43.11–00.37.04.08]

And it just makes sense: communication using standard language, so that all the people involved in the incident are on the same page. What does this term mean? We don’t want to get terms crossed. And so what it’s trying to do is standardize those.

Matt:

[00.37.04.09–00.37.17.09]

So have you been in scenarios where someone had not thought through that part, as far as, they’re not actually in the room with us, how do we dial them in?

Chris:

[00.37.17.10–00.37.25.16]

Oh yeah. In the five or six years since the pandemic, we’ve done so many OT incident responses remotely.

Matt:

[00.37.25.17–00.37.26.06]

Yeah.

Chris:

[00.37.26.07–00.37.59.01]

My team, I might be remote, but we have somebody that’s going on site. Or maybe they have a contract with their local engineering firm or automation firm, and they’re contracted to be there and have hands on keyboards. There may be a union, there may not be; there are a lot of different situations there. But it’s all about communication, virtually or on the phone or in person or a combination of all three.

Chris:

[00.37.59.07–00.38.22.22]

All those things are key. And then obviously there are back channels. You have to have back channels that are not on the compromised network. So if your email is compromised, you have to switch to a backup set of emails, or a Signal channel, or some other previously agreed upon channel. Or even in the heat of the moment, if they don’t have a backup channel,

Chris:

[00.38.23.00–00.38.25.00]

those get stood up.

Matt:

[00.38.25.02–00.38.45.18]

Regardless, one way or the other. Yeah, absolutely. I think it was one of the Dragos executives who told us one time that when they do these OT tabletops, three times out of four someone hasn’t even thought of what happens if we can’t just use our Teams chat or our Slack or whatever.

Chris:

[00.38.45.20–00.39.09.06]

Well, I ask them if they have it in their business continuity: what happens when you lose power? We had radios at the power company, right? And then if those radios went down, we had satellite phones in different offices as the backup to the backup. Primary is the cell phone, and all of a sudden backup is the radios.

Chris:

[00.39.09.10–00.39.25.02]

Backup to that is the satellite phone. So some of the sites may have those, some of them may not. They may not have thought about it. But that’s what tabletops are there for: to ask the question and pressure test what you would do in that scenario.

Matt:

[00.39.25.03–00.39.51.22]

Absolutely. I want to go back to a question off of a question we talked about earlier, when it comes to some of the incompatibility with the standard cyber tools and things like that. I forget where I saw something you were writing, or maybe it was a speaking event or something.

Matt:

[00.39.51.23–00.40.20.03]

You were talking about cyber-informed engineering: the idea of anticipating cybersecurity needs as you’re building these physical devices and tools, so that, one, they’re more resilient, but also things are a little bit more standardized. For listeners who haven’t really thought through, or aren’t aware of, cyber-informed engineering, what is it?

Matt:

[00.40.20.05–00.40.29.15]

What does it mean? Because I’m sure my take was a bit too high-level there. From your perspective, what does it mean to treat cyber from that perspective?

Chris:

[00.40.29.17–00.40.55.14]

Right. So I can give it from an engineering perspective, when I was in college. I’m an electrical engineer. And so when I was in school, in the late 90s, early 2000s, and I graduated in 2002, there was no security in my engineering classes. We were not taught cybersecurity. We had a computer class where we would run into the folks that were into hacking, and, really cool thing I could do.

Chris:

[00.40.55.14–00.41.28.06]

And that was it. We didn’t have any of that. And so the body of knowledge of engineering did not have security. But now, fast forward 20, 25 years, we have cyber-informed engineering, which seeks to bake into the engineering design process thinking out the consequences of any impact in the design, any failure modes, and then working backwards.

Chris:

[00.41.28.06–00.42.02.20]

Could we map these to cyber issues or designs? And then how do we therefore make it more resilient, make it less susceptible, or even eliminate it completely as a possibility? So that’s my high-level take on it. The folks that I know at the Department of Energy, and all those who have worked on CIE, Andy Bochman and all these others, Tim Roxey, yeah, I love that guy.

Chris:

[00.42.02.21–00.42.43.04]

And Ginger, she’s so incredibly smart. All that knowledge base that’s out there on it, people can find the thing that speaks to them and then get plugged into how that works, whether you’re an engineer, a network person, or an OT security person. So it’s something for everyone. But really, it’s designed for helping protect when you’re building something new or adding something, and doing the whole thought process around informing yourself of the cybersecurity risks and designing that into your system.

Matt:

[00.42.43.06–00.43.12.04]

So, CISA, and over the past year or two there have been a number of government-sponsored analyses of OT risks from a cyber perspective of engineering and design. There seems to be a theme that I’ve picked up on, where most of the paper is dedicated to a number of things.

Matt:

[00.43.12.04–00.43.41.00]

But recently I’m noticing, and it’s always a big deal when governments call out vendors, they seem to keep harping on certain things that are industry standard outside of the OT side and don’t seem to be baked into the design yet. They’re talking about 2FA and MFA capabilities. That’s one that I see a lot.

Matt:

[00.43.41.01–00.44.09.09]

But it seems like OT, because it has benefited from being in the shadows from a vendor standpoint, that obscurity and that air-gapped nature has allowed, in some respects, at least from my outside perspective, these vendors to get a little behind the curve versus some of the engineering standards outside of it.

Matt:

[00.44.09.09–00.44.28.12]

Because, well, we’re just air gapped, no one, it’s all proprietary, nobody needs to know what’s happening in here. But it feels like now we’re all blind, because we can’t see what’s going on in here. Do you see cyber-informed engineering impacting that, or do you think that’s a problem at all?

Chris:

[00.44.28.14–00.45.08.13]

I don’t think it’s a problem, because the customers that buy those things have their requirements for it. So you’re still going to have insecure OT protocols that have no encryption, because physics doesn’t allow for the encryption. For instance, if you tried to encrypt, say, a circuit breaker with DNP3: the way you’re designed today, you’re sending a trip command to your substation relay and it’s sending it over the network for some reason, instead of being hardware. Most of the time it’s hardware.

Chris:

[00.45.08.13–00.45.36.10]

But let’s just hypothetically say you cannot encrypt it, because it will cause a wire to burn down. If you have 100 milliseconds to clear a fault before the wire starts to melt, then depending on the speed and depending on different constraints, encryption in that path could easily cause a physical problem.

Matt:

[00.45.36.11–00.45.37.13]

Yeah.

Chris:

[00.45.37.14–00.45.46.08]

Sometimes you need milliseconds to do an action. And so physics dictates things like that.

Matt:

[00.45.46.08–00.45.46.23]

Sure.

Chris:

[00.45.47.03–00.46.25.17]

And so then there are other things, like hardware two-factor, where you really need to have the horsepower in these devices. Nowadays there’s more chip power, so you can do these things. But a lot of it’s not necessary if you have physical gates and locks and physical two-factor access to that device. That can still be two-factor access, if you don’t allow remote access to it. If you allow remote access to it,

Chris:

[00.46.25.17–00.46.52.19]

yeah, you need to use remote two-factor through a standard. There are so many different things like that that won’t be in a PLC, because, one, it’s not asked for by the asset owner. They’re not going to ask for secure Modbus, probably, in most cases. And then there’s, okay, now we need someone to manage it, and manage all the keys and the encryption and all that stuff.

Chris:

[00.46.52.20–00.47.20.00]

So it adds another layer of complexity, even though it is adding security. So you can’t just willy-nilly throw in security and make things more secure without having a plan. So I don’t think the blame is on the vendors. There are obviously vendors out there that have security features, that can put in secure Modbus and secure firmware downloads to these systems.

Chris:

[00.47.20.00–00.47.43.16]

There’s been a lot of improvement since Stuxnet. So I don’t think the vendors are the CIE issue. I think the context in which that whole system is built is the whole purpose of CIE. So having a secure device, or capabilities in that PLC, for instance, is one component of the whole burrito.

Matt:

[00.47.43.16–00.48.06.09]

Makes sense. I know we have some constraints on time today. There’s one question we always wrap up with here on the Lounge, but before I jump to that: any closing thoughts or comments from your side?

Chris:

[00.48.06.11–00.48.32.17]

Thanks for having me on, first of all, Matt. And anyone listening: you can be involved in OT security, whether it’s your job or not. Be aware of what’s out there. Keep your stuff off the internet, and believe in yourself. You can help with this. And it starts with just asking yourself a question, and doing the basics, doing them well.

Chris:

[00.48.32.19–00.48.49.21]

Read some of those findings from 20 years ago, 15 years ago. The first Stuxnet alert from back then: some of those recommendations still work today. So don’t forget about those things, and be curious.

Matt:

[00.48.50.02–00.49.15.22]

Absolutely. Stay curious. All right. So the question is, imagine you’re at the bar, and everybody has a bar in mind when they think of this: the type of bar that’s nice enough where you don’t have to yell to be heard. You’re at one end, and at the other end is somebody from the security industry that you’ve always wanted to have a drink with, talk to, or ask a question of. And you being a veteran,

Matt:

[00.49.15.22–00.49.31.11]

that’s probably a pretty short list, I would imagine. There’s probably not a ton of people. But that person is there, and I understand you’re a bourbon guy. So what are you ordering for yourself, and who’s at the other end of the bar?

Chris:

[00.49.31.13–00.49.36.11]

All right, so I always order a Sazerac. That’s my test on how good a bartender is.

Matt:

[00.49.36.11–00.49.37.07]

Nice. Okay.

Chris:

[00.49.37.13–00.49.57.05]

And either they know how to make it and they make it for me, or they don’t have the ingredients but they know how to make it, so that’s still cool. They may not have the absinthe or whatever. And the last one is, what’s this? You want a Bud Light? I’m like, I’m out. On to the next bar.

Chris:

[00.49.57.07–00.50.11.15]

So, who would be sitting across from me: I would love to sit down with Cliff Stoll. I got to meet him at DEF CON and it was one of my favorite things ever. His book, and his excitement, he’s…

Matt:

[00.50.11.16–00.50.13.06]

For folks not familiar.

Chris:

[00.50.13.07–00.50.39.16]

Well, he had one of the first discoveries of a network intrusion. He’s probably the godfather of network forensics and network security monitoring, before it was invented. But in 1986 they caught a hacker in Germany. And if you go find my LinkedIn post, I posted my picture with him, and then the hacker, Markus, replied in the comments, which just blew my mind.

Chris:

[00.50.39.16–00.50.46.22]

I’m like, oh my gosh. And look, if Cliff doesn’t drink, not a problem. We can drink water. I’d just love to have…

Matt:

[00.50.46.23–00.50.48.11]

Soda and lime.

Chris:

[00.50.48.13–00.50.53.14]

I could talk to somebody with that much enthusiasm all day. That would make my day.

Matt:

[00.50.53.15–00.50.55.04]

Yeah, yeah.

Chris:

[00.50.55.06–00.50.55.19]

That’s great.

Matt:

[00.50.55.20–00.51.02.20]

Well, Chris, this has been really, really good. Thank you for the time here.

Chris:

[00.51.02.22–00.51.05.17]

Thank you for having me. And you all be good.

Matt:

[00.51.05.18–00.51.22.10]

Absolutely. And to our listeners, thanks for taking the time to dive in with us here today. You can always find us over at armortext.com, or at lounge@armortext.com. And until next time: be well, stay curious, and do good work.

Matt:

[00.51.22.12–00.51.55.12]

We really hope you enjoyed this episode of The Lock and Key Lounge. If you’re a cybersecurity expert, or you have a unique insight or point of view on the topic — and we know you do — we’d love to hear from you. Please email us at lounge@armortext.com, or our website: armortext.com/podcasts. I’m Matt Calligan, Director of Revenue Operations here at ArmorText, inviting you back here next time, where you’ll get live, unenciphered, unfiltered, stirred, never-shaken insights into the latest cybersecurity concepts.

Search