Thursday, May 27, 2021

A SOC Tried To Detect Threats in the Cloud … Your Won’t Believe What Happened Next


Now, we all agree that various cloud technologies such as SaaS SIEM help your Security Operations Center (SOC). However, there’s also a need to talk about how traditional SOCs are challenged by the need to monitor cloud computing environments for threats. In this post, I wanted to quickly touch on this very topic and refresh some past analysis of this (and perhaps reminisce on how sad things were in 2012).

Back in my analyst days, I’ve noticed that some traditional organizations tried to include their cloud environments in the scope of their security monitoring at some point in their cloud migration journeys. Surprisingly (Hey … you surprised about it? No? Thought so!), some of these projects have not gone well. SOC teams were not equipped to deal with various cloud challenges (old paper on this). There were also cases where both business and IT migrated to the cloud, but security was left behind and had to approach cloud challenges with on-premise tools and practices. Essentially, security was left behind … again.

Here, we wanted to quickly summarize some of the challenges, covering the usual range of people, tools, and processes:

  • Uncommon log collection methods (compared to on-premise systems). Cloud providers haven’t necessarily simplified this journey for customers, even though, compared to 2012, decent logs actually exist today in many cases.
  • Telemetry data volumes may be high (especially from all those web-facing production systems); this has sometimes led to “log fragmentation” where cloud logs never make it to a SIEM, but are left to rot in some storage buckets in the cloud.
  • Egress costs are there sometimes, especially if you want to move the logs from one cloud to another for analysis.
  • Alien licensing models for security tools (compared to on-premise), some teams can’t afford what they used to be able to afford on-premise or they can’t afford a new cloud-native tool in addition to the on-premise tool they already have.
  • Alien detection context— instances, containers, microservices, etc — has confused many teams born and raised on server names and IP addresses for context. This topic is big enough to be explored in a dedicated post later.
  • Lack of clarity on cloud detection use cases is there despite useful resources like ATT&CK Cloud. Sadly, cloud providers haven’t necessarily simplified this journey for customers either, and many traditional SOC teams are not sure what to detect in the environments that their business is using today (“is this container access bad?”).
  • Also, there is a lot of cloud; this means governance sprawl causes visibility gaps for the SOC. Examples include shadow IT (“BYOCloud” and SaaS purchased by departments) as well as other cloud sprawl (that is why people are reaching for all those novel attack surface management tools; this should help).
  • SOC teams lacking cloud skill in general; complex public/hybrid/multi — cloud scenarios require more extensive knowledge of various technologies, their security implications, diverse (and alien) data sources, while SOC teams are too busy doing D&R to grow their cloud skills.
  • For those organizations trying to stick to old on-premise tools many other challenges abound; tools don’t support many cloud telemetry sources — they lack collection machinery, parsing/analysis, use cases, useful visuals, etc. Also, log support is often not done at “cloud speed.”
  • Lack of input from SOCs into cloud decisions, ranging from provider choices to IT architecture (and even security architecture). Frankly, many SOC teams are too busy and too focused on threats and don’t have a dedicated headcount focused on preparing their organization for the cloud change …

Huge thanks to Iman Ghanizada (“the Certs Guy”) for his contributions to this post.

Related posts:


Originally published at Medium.

A SOC Tried To Detect Threats in the Cloud … Your Won’t Believe What Happened Next


Now, we all agree that various cloud technologies such as SaaS SIEM help your Security Operations Center (SOC). However, there’s also a need to talk about how traditional SOCs are challenged by the need to monitor cloud computing environments for threats. In this post, I wanted to quickly touch on this very topic and refresh some past analysis of this (and perhaps reminisce on how sad things were in 2012).

Back in my analyst days, I’ve noticed that some traditional organizations tried to include their cloud environments in the scope of their security monitoring at some point in their cloud migration journeys. Surprisingly (Hey … you surprised about it? No? Thought so!), some of these projects have not gone well. SOC teams were not equipped to deal with various cloud challenges (old paper on this). There were also cases where both business and IT migrated to the cloud, but security was left behind and had to approach cloud challenges with on-premise tools and practices. Essentially, security was left behind … again.

Here, we wanted to quickly summarize some of the challenges, covering the usual range of people, tools, and processes:

  • Uncommon log collection methods (compared to on-premise systems). Cloud providers haven’t necessarily simplified this journey for customers, even though, compared to 2012, decent logs actually exist today in many cases.
  • Telemetry data volumes may be high (especially from all those web-facing production systems); this has sometimes led to “log fragmentation” where cloud logs never make it to a SIEM, but are left to rot in some storage buckets in the cloud.
  • Egress costs are there sometimes, especially if you want to move the logs from one cloud to another for analysis.
  • Alien licensing models for security tools (compared to on-premise), some teams can’t afford what they used to be able to afford on-premise or they can’t afford a new cloud-native tool in addition to the on-premise tool they already have.
  • Alien detection context— instances, containers, microservices, etc — has confused many teams born and raised on server names and IP addresses for context. This topic is big enough to be explored in a dedicated post later.
  • Lack of clarity on cloud detection use cases is there despite useful resources like ATT&CK Cloud. Sadly, cloud providers haven’t necessarily simplified this journey for customers either, and many traditional SOC teams are not sure what to detect in the environments that their business is using today (“is this container access bad?”).
  • Also, there is a lot of cloud; this means governance sprawl causes visibility gaps for the SOC. Examples include shadow IT (“BYOCloud” and SaaS purchased by departments) as well as other cloud sprawl (that is why people are reaching for all those novel attack surface management tools; this should help).
  • SOC teams lacking cloud skill in general; complex public/hybrid/multi — cloud scenarios require more extensive knowledge of various technologies, their security implications, diverse (and alien) data sources, while SOC teams are too busy doing D&R to grow their cloud skills.
  • For those organizations trying to stick to old on-premise tools many other challenges abound; tools don’t support many cloud telemetry sources — they lack collection machinery, parsing/analysis, use cases, useful visuals, etc. Also, log support is often not done at “cloud speed.”
  • Lack of input from SOCs into cloud decisions, ranging from provider choices to IT architecture (and even security architecture). Frankly, many SOC teams are too busy and too focused on threats and don’t have a dedicated headcount focused on preparing their organization for the cloud change …

Huge thanks to Iman Ghanizada (“the Certs Guy”) for his contributions to this post.

Related posts:


Originally published at Medium.

Thursday, May 13, 2021

SOC Trends ISACA Webinar Q&A

A few days ago we did a very well-attended webinar focused on the modern Security Operations Center (SOC) approach (see “Trend for the…


A few days ago we did a very well-attended webinar focused on the modern Security Operations Center (SOC) approach (see “Trend for the Modern SOC” for a replay link). We got a lot of great questions, and just like in the good old times, I am writing a blog where I cover some of the answers.

Q: You mentioned that SOC is first a team: which skills are expected to distinguish the “basic” SOC from the modern SOC?

A: From our presentation, it’s relatively clear that such skills include threat hunting, threat intelligence, data analytics, and others. These are less common at traditional SOCs, but they power the enabling capabilities of the modern SOC that we discussed. Also see this paper.

Q: Can we achieve a fully automated AI/AL based — OODA? Fully automated onboard log sources, threat detection rule creation, playbook creation, response, automated integration, and execute.

A: Today and in the near future, I do not believe that a complete automation of most SOC processes is possible — see link. Frankly, the most troublesome part is at the end of the chain where automated response and other actions happen. They’re also other situations that require human decision-making to deal with a high degree of uncertainty. Finally, even onboarding for many tricky telemetry sources requires humans to iterate and tweak the configurations sometimes.

Today automation is more widespread in the areas like detection (create alerts) and triage (enrich and confirm alerts), but a lot less widespread in remediation and data onboarding. I do not expect any massive change here soon, but as organizations adopt more public cloud, automation will grow in these areas too. So, ask me again in, say, 5 years, but really more like 10.

Q: What is the difference between SIEM and SOC?

A: This distinction should be abundantly clear from the presentation : SIEM is a particular security tool while SOC is the name of a team together with associated processes and tools they use (including, for many SOCs, a SIEM).

That is why I am always a bit skeptical when I hear SOC-as-a-service, and I prefer the term MDR instead.

Q: If we can’t get top tier attackers out of our network — how does a company handle that risk?

A: It is hard to give prescriptive advice here as this is a tough challenge and it falls into the heavy “it depends” territory. Most organizations that encountered this will need to call for help and invite a 3rd party incident response team to help them investigate and ultimately get the attackers out.

As sad as it may sound, it is entirely possible that you will encounter an attacker who is just better than your own team (even if your team is good). In this case you will need also to ask for help and there is no way around it, cost notwithstanding.

Q: Am I correct in understanding that we are hearing advocacy of taking a risk-based approach to design and management of a modern SOC? I think that is what I am hearing here. Correct?

A: It’s not entirely clear what “risk-based” means here in your question. Most currently functioning SOCs are not exactly built based on a fixed checklist from a compliance regulation. In that sense, most SOCs I encountered are at least somewhat risk-based.

Q: Can you touch on dispersed SOC staff especially in a COVID environment. Is it practical to spread your staff remotely across the USA? Outside USA?

A: Follow the sun model for SOC is very well known and many global organizations practice exactly that, even if they do so with distributed teams, not people. However, it is also very clear that during the current pandemic many detection teams and formal SOCs operated in a distributed manner. I feel that the jury is still out regarding whether they were more or less productive, but for sure it was not a failure, hence the model may work.

Q: What makes a good SOC?

A: Frankly, I do not think I have a short answer to this question (long answer, another). I think a bad SOC is the one that over indexes on technology and had excessively rigid processes, while a good SOC is the one that really focuses on people, and then on process/workflow.

Q: Regarding SOC tools, what do you think about AI tools used in SOC?

A: This is of course a fascinating question that I spent a good number of years trying to answer, starting from the time I was an analyst. I think over time I’ve reached a position that the only way is to be skeptical about AI for security in the short term, but ultimately optimistic in the long term.

Naturally, we have a lot of vendors with madly (sorry, no links here …) overblown claims about how their ML/AI tools help security analysts. However, just as AI evolves to help other areas of human endeavor, cyber security is not an exception.

Today, the most likely machine learning — based tool that you will encounter in a SOC is some form of anomaly detection such as a UEBA tool or an NDR. Of course, these tools work and they produce alerts that are often useful (just as regular rule-based alerts). However, it’s very clear to me that today there’s no magic of “cyber AI” in today’s SOCs.

Q: What skill sets do you look for in threat hunter personnel?

A: This is a madly difficult question to answer, and I did try to answer it in my analyst days. Given that great hunting is ultimately an art, but that said artist need to also be a top-tier technologist, defining a skill set is very difficult. For sure, threat knowledge, deep IT technical knowledge and creative thinking are all must-haves here.

Q: How can a small company/start up weigh out talent vs tools and cost?

A: This is yet another question that I’ve tackled back in my analyst days. It became very clear to us early on that as smaller organizations will use more third-party services, what some would call outsourcing. Some won’t have a SOC at all and will utilize an MSSP or an MDR provider. Others will use a hybrid model.

Naturally, this comes with its own pitfalls and benefits. The one key pitfall is that one can’t assume that you can pay somebody money and they take security off your hands…

Q: What about SOCs as a Service and Internal SOCs? Would your recommendations apply for both?

A: If by SOC-as-a-service, you mean using an MSSP or an MDR provider, then some of the recommendations from the webinar do apply as well. An MSSP provider may follow a more traditional SOC approach or they may use the modern SOC elements discussed here. Many MDR providers I encountered practice modern SOC approach.

Q: What is the difference between Security Operation Center and Security Operation Control?

A: I have never encountered an industry term “Security Operation Control.” I don’t know what that is. Google (the search, that is) does not seem to, either.


Originally published at Medium.

Wednesday, May 05, 2021

Not the Final Answer on NDR in the Cloud …


Back in my analyst years, I rather liked the concept of NDR or Network Detection and Response. And, despite having invented the acronym EDR, I was raised on with NSM and tcpdump way before that. Hence, even though we may still live in an endpoint security era, the need for network data analysis has not vanished.

As we discussed during this recent webinar, this is not about competing with endpoint or endlessly arguing about what security telemetry is “better.” This is about reminding the security leaders and technologists that network telemetry matters today! Not only in the 1980s (when tcpdump was born), 1990s, 2000s, 2010s, but today in 2020s.

To summarize, network security monitoring still matters because you can monitor unmanaged devices (BYOD, IoT, ICS, etc.), detect threats with no agents, offer broad coverage from a few points, and be out of band (go and see my old Gartner paper for details).

Still, I see a few common misconceptions (more details here in this webinar) about network security telemetry data. I wanted to cover a few and then focus on ONE, in particular.

  • You cannot monitor encrypted data: as I discussed here, encryption for sure saps some of the value of network security monitoring, but it does not destroy it. Both layer 3 (flow) and layer 7 (rich metadata) observation have value for encrypted data whereas full packet capture perhaps does not.
  • Network monitoring is only an auxiliary control, you need endpoint first: Well, OK, maybe, but so what? You may need an endpoint first,I’ve seen enough environments where it’s the truth. The point is that you need an endpoint first, but then you need NDR to cover the gaps, unmanaged devices, etc, etc.
  • “PCAP or it didn’t happen”: many years ago, before we had Bro/Zeek and the choices were “flow or pcap”, this may have been true. But you know what? In 2021, you are not saving full packet captures for weeks or months. Perhaps we have to change the slogan to “zeek decodes or it didn’t happen”?
  • Network traffic is too expensive to capture: this is not a misconception at all, if you see full packet capture as the way to go. It would be prohibitively expensive in most modern environments. However, you can get a lot of value from rich L7 metadata and this is much less expensive (but also more useful than mere flows)
  • Network data is not helpful in the cloud: while comparatively fewer people capture and monitor traffic in the cloud, the interest to do this grows rapidly. This is also discussed in depth below

Let’s now drill down into the last point:

Why some people think that NDR in the cloud is an anti-pattern (I prefer the term “worst practice” instead)?

  • In the cloud everything is locked down and immutable, what’s the point of traffic capturing? — Sure, but it is really? In theory, it should be, but is it in your cloud?
  • Everything is encrypted, so what’s there to sniff? — We already addressed this above in general for both cloud and on-premise. NDR has value for encrypted networks.
  • Cloud logs and this new fancy observability stuff provide visibility, why sniff traffic? — Well, are these logs complete and available, and can be leveraged for security value? Sometimes the answer is “yes”…
  • I can do flows logs in the cloud, I don’t need “costly” packets. — Same as on-premise, flow logs may not do the trick for the threat detection needs you have.
  • Applications are dynamic and everything changes so captures become useless over time. — This just works that an NDR vendor needs to work harder, but not that NDR is not useful.

Go and see this webinar for additional discussion.

On the other hand, I’d like to say that NDR fits well with the public cloud today.

  • Your main on-premise tool — EDR — may not be available at all (containers, etc)
  • Some cloud architectures do use what on-premise would be called a flat network, hence NDR is very useful for East/West visibility.
  • Cloud API logs are not exhaustive, but they are voluminous, often have inconsistent schemas and sometimes not designed for security use cases.
  • In-app observability for security is not common yet, even if it is coming.

Thus, NDR lives on in the cloud!

Now, if you live mostly in SaaS applications, NDR approach may not be a fit, but I have not seen many large organizations with no IT whatsoever, just SaaS management (I bet they do exist, but they are not common).

For everybody else, in the cloud or not, NDR works. This applies to both virtual machine environments and modern cloud environments, even if not equally …


Originally published at Medium.

Dr Anton Chuvakin