Thursday, December 22, 2022

Cloud Security Podcast — Two Years Later or Our Year-End Reflections for 2022!

We have been running our Cloud Security Podcast by Google for almost 2 years (TWO YEARS!) and since we are on a break now, I wanted to…


We have been running our Cloud Security Podcast by Google for almost 2 years (TWO YEARS!) and since we are on a break now, I wanted to reflect a bit, while Tim is relaxing on a beach somewhere warm and “hammy” 🙂

So, we aired 102 episodes, but what was new in 2022?

Here is how 2022 word cloud of episode titles looks like

(src)

What to expect from us in 2023?

  • A weekly podcast episode and a few specials here and there — of course!!!
  • Perhaps a live video of our recording session — that will be fun! Should we post the audio to YouTube, BTW?
  • Some stuff that is coming in Q1 2023 includes episodes on BeyondProd, our security guardrail magic, security architecture (with more cloud migration challenges!) and a curious episode on our approach to vulnerability management. More “CISO meets cloud” episodes are planned as well!
  • The 2023 season opener on January 9 would be epic as well ... because Mandiant!

Now, let’s look at our basic metrics:

Top 5 episodes overall so far:

  1. “Confidentially Speaking“ (ep1)
  2. “Data Security in the Cloud“ (ep2)
  3. “Megatrends, Macro-changes, Microservices, Oh My! Changes in 2022 and Beyond in Cloud Security” (ep47)
  4. “Automate and/or Die” (ep3)
  5. “How We Scale Detection and Response at Google: Automation, Metrics, Toil“ (ep75)

Share your favorites in comments or on social media (LinkedIn, Twitter)?

We also understand that 102 episodes is a lot and that we cover many topics in a wide field of cloud security. So, if you want to cherry pick, here are some fun playlists (on Spotify):

So, a call to action:

P.S. Of course we asked the future AI overlords about what we should do next, here are the answers.

And of course, we are us, so here is what we did next:

:-)

Related:


Originally published at Medium.

Thursday, December 15, 2022

Combined SOC Webinar Q&A: From EDR to ITDR and ASO … and ChatGPT

In recent weeks, I did two fun webinars related to Security Operations, and there was a lot of fun Q&A. The questions below are sometimes…


In recent weeks, I did two fun webinars related to Security Operations, and there was a lot of fun Q&A. The questions below are sometimes slighting edited for clarity, typos, etc.

For extra fun, I had ChatGPT answer some of them, to see if it can replace me :-)

So, first, ISACA webinar “Modernize Your SOC for the Future” focused on our Autonomic Security Operations vision.

Q: If not called SOC, would you like to share what you have named the team? [this question is related to the fact that at Google, there is no team called “SOC”]

A: The discussion about naming the security operation center comes from the longer debate about whether SOC includes just the analysts watching the screens or the infrastructure and processes for producing the alerts, threat research, detection creation, etc. As a result, some organizations assume that SOC stands only for a team content watching the screens and they do not want to call their combined/integrated team the same name

The names I encountered are just detection and response team, D&R team, etc with some name choices unique to Google. Please don’t say “XDR team”, if at all possible :-)

Q: The “modern SOC” looks great on paper, but it’s harder to find analysts who have necessary skills to focus on threats when you have to deal with attrition. What can you propose to find skilled people for these roles faster?

A: Indeed, the challenges with using the analysts for creating detection content and pursuing threats implies that they have the skills to study the threats and to create detection content. Some people assume that this problem is solved primarily by hiring, but in our opinion it is solved by motivation, training, retention, and, perhaps last, by hiring.

We are certainly not suggesting you fire your SOC analysts and hire some mythical full- stack security engineers instead (or robots, for that matter). What you do is start to select, train and motivate your analysts to explore outside of passing the alerts around. More details on this are provided in the original ASO paper.

Q: Could you please explain a bit more on the use case library?

A: When we refer to the use case library in the context of SOC, we mean a collection of your rules, playbooks and other detection content, with its associated processes. Think about it, a book library is a collection of content for people to read while a use case library is a collection of use case content for the detection tools to run. Such a library comes with its own processing workflows such as use case creation, tuning and modification, and retirement.

Details of this can be found in this “How to Create and Maintain Security Monitoring Use Cases for Your SIEM” Gartner paper (sorry for the paywall, analysts need to eat)

Q: Please expand Threat Hunting with examples, any risks?

A: I would defer to stuff that others have written and to some of my own writing from the past to define threat hunting. Here I would say that merely searching for indicators may be part of hunting but it isn’t the entire thing, for sure.

To me, the more interesting part of your question is a question about risks of threat hunting. During the webinar, my response focused on the fact that while there are no risks of hunting, there may be associated risks of uncovering something that you don’t know how to deal with. While we can joke all we want about calling for help in this case (“call Mandiant!”), that is absolutely the correct advice in this situation.

See “My “How to Hunt for Security Threats” Paper Published” and see also “Threat Hunting Is Not for Everyone.” And also “Beware: Clown-grade SOCs Still Abound” for advice on when NOT to hunt.

P.S. And if you don’t trust humans, here is the answer from your friendly neighborhood robot, ChatGPT by OpenAI

Q: What are the baselines of SOC /SIEM Implementation? Is there any standard for SIEM/SOC?

A: These are essentially four questions, each fairly complicated. They’re definitely industry guidance documents on implementing SIEM (google for it, I don’t have a shortlist of faves somehow), and I have written a fair share of them while at Gartner. However, the differences between organizations drive differences in how they implement SIEM tooling and run their SOC teams. Probably the closest to the standard SOC guidance is the “11 Strategies of a World-Class Cybersecurity Operations Center” book, that has recently been updated

Beyond that and especially if you cannot access Gartner content, Google is your friend.

Q: Should SOC staff be made up of a multi-lingual team to handle the global diversity of threats?

A: My answer to this question will be determined by whether you treat your threat intelligence team as a part of the SOC. At other places, the threat intelligence team is a peer to the SOC rather than a part of it. If you are planning to stand up a large research team, you probably would need multilingual talent. However, for a more traditional SOC that seems excessive.

See “About The Tri-Team Model of SOC, CIRT, “Threat Something” for details.

Q: Do you believe in interdepartmental training examples from Business Analyst to SOC analyst?

A: I believe that literally any background may lead to somebody being a great security professional, and can perhaps point at examples of each. So, definitely a business analyst can become a SOC analyst.

In other regards, this is a difficult question as a lot would depend on the person. Sure, some best security analysts in particular and security professionals in general come from all sorts of backgrounds, including some that are very unusual — like music or physics. However, such a person will definitely have to learn a lot about information security, technology and practice, etc and cover both technical and non-technical domains of security.

Q: What’s the most common problem encountered when you have SOC managed by a third party vendor?

A: This is something that I’ve mentioned many times in my previous analyst writing. The most common problem between a client and the managed services provider is an expectation mismatch. When clients approach a third party for managed security services and their expectation is that “they would pay money” and “they would get security,” disaster is almost certain. A more healthy model that solves this problem is thinking of your work with such third party as jointly operating (“JointOps” anybody?) your SOC, rather than using the dreaded “O word” — outsourcing.

Let’s see what the AI thinks:

Now, the questions below are from the BlackHat webinar “SOC Modernization: Where Do We Go From Here?”

Q: With agile application development in a large organization — are there tips to better integrate those native cloud apps, pipelines, etc. with the SOC?

A: I have an entire presentation largely on that and while it does not go into all the details, I would point you there. Definitely more work on this is needed, as I see many who struggle with “fast DevOps, meet slow Security” in various forms of this disease.

Q: Do you think that “SOC” includes development resources or is it a more investigatory function or … any recommendation for small and medium sized companies with limited “SOC” resources?

A: This topic comes up a lot and I would say that there are companies that treat their SOC as only the team watching the screens while they have the more engineering components outside of the SOC. Perhaps you can call it a more traditional model.

However, in my opinion, a modern SOC tightly integrates security analysts with people who develop detections. Ultimately, this model evolves more closely to our ASO-style operation where they are the same people

Q: With the rise of SOAR/XDR/etc. and other tools, do you still see the SIEM as the main engine or operating system of the SOC?

A: This is a fascinating question that has a lot of nuance in response. The short answer is that I still see some log analysis capability (such as SIEM) to be the center of many SOC teams.

However, would I insist on this as before? No, at this point I have seen enough of the EDR-centric SOC teams (for example) that actually use their endpoint tools as a primary console and they treat log analysis as an auxiliary.

I’ve also seen organizations that center their SOC on their SOAR tool which then access data from SIEM or EDR, but ultimately the analysts live inside the SOAR console most of the time.

In the future, this may well be rebalanced. It is very possible that we are at peak EDR and as cloud native services are adopted more widely, the importance of endpoint would again decrease yet the importance of logs — and this likely means SIEM — will increase.

See “Can You Do a SIEM-less SOC?” for more details.

Q: In the cloud and on-premises mixed environment, do you think installing the user behavior analysis agent on the laptop makes sense?

A: In my experience, endpoint-based employee monitoring addresses special use cases that some organizations may have, while many don’t. So I would say that endpoint-based user monitoring tools remain popular at select organizations such as those that place high importance at insider threats. For them, installing the agents for deeper employee monitoring makes sense for others, I’m not so sure…

Q: Given your comment on EDR-centric SOCs, do you think SOCs should migrate to be more Identity-centric as customers migrate to the cloud / SaaS services?

A: My former colleagues at Gartner just coined the term ITDR that stands for identity thread detection and response. While we can debate whether a new acronym is needed here, this is not debatable: cloud environments place increased importance on this type of monitoring and detection.

However, a lot of identity centric monitoring is in fact a very traditional feature of SIEM and UEBA (now part of SIEM) going back almost 20 years

Q: Is the major trend here using a premises-based SOC to monitor/manage more cloud resources, or pushing more SOC functions to the cloud?

A: Well, “Today, You Really Want a SaaS SIEM!” — so I very much believe in most SOC technologies being cloud-backed and cloud-native. As we say in our original ASO paper, deploying such tools as cloud-based SIM and cloud-based EDR becomes almost the only choice for many organizations. If I have a limited team, I’d rather this team use the tools and deliver security value rather than maintain and manage the tool.

Enjoy!


Originally published at Medium.

Monday, November 21, 2022

Security Incident Response in the Cloud: A Few Ideas

This quick blog is essentially a summary of our (joint with Marshall from Mandiant) Google Cloud Next 2022 conference presentation (video)…


This quick blog is essentially a summary of our (joint with Marshall from Mandiant) Google Cloud Next 2022 conference presentation (video) and a pointer to a just-released podcast on the same topic — security incident response (IR) in public cloud.

In our Next presentation, we only had 18.5 minutes to present a few fun and insightful things about security incident response in the cloud.

Here’s what we decided. We focused on three challenges that we observed with organizations preparing for security incident response in the cloud, these are:

  • Skills: Cloud IR requires both solid security incident response skills and equally solid cloud native technologies skills
  • Joint nature: Many (but not all) cloud incidents will involve a CSP, 
    and many will involve a client, a cloud provider and one or more security service providers
  • Data: In many cases, telemetry, logs, traces data won’t be available or won’t be available via familiar mechanisms.

Next, we decided to focus on critical differences as well as similarities — no less critical — between security incident response on premise and in public cloud.

Now, if you want a one — line summary, the similarities mostly stem from the facts that the threat actors ultimately need to achieve their goals, and that responders need to know the environment to respond well (Duh, no brainer? Perhaps, but it affects how you do IR).

Here are the similarities:

  • Data preservation requirements.
  • Comprehensive understanding of the environment.
  • Standard investigative techniques have not changed.
  • Log data needs to be retained, normalized, and analyzed. Time Zones and time skew must be addressed.
  • Each incident is different.

Similarly, the differences mostly stem from the fact that cloud technology is often different, and operational practices for the teams behind called environments are different as well. As a side note, while people want to focus on logs from cloud services and containers, the fact that environments are just run differently, and jointly with your cloud provider partner, and that affects IR quite a lot.

Here are the differences:

  • Ephemeral & dynamic nature of the cloud.
  • Deep technical expertise of cloud native services required.
  • Different baseline and norms.
  • Log data retention, understanding, context and volume
  • Reliant on CSP & customer for relevant data.

Please watch the video and listen to the podcast. By the way, they cover completely different things, and in the podcast specifically, we share some deep secrets of how Google does IR in the cloud …

Related blogs on cloud security:


Originally published at Medium.

Security Incident Response in the Cloud: A Few Ideas

This quick blog is essentially a summary of our (joint with Marshall from Mandiant) Google Cloud Next 2022 conference presentation (video)…


This quick blog is essentially a summary of our (joint with Marshall from Mandiant) Google Cloud Next 2022 conference presentation (video) and a pointer to a just-released podcast on the same topic — security incident response (IR) in public cloud.

In our Next presentation, we only had 18.5 minutes to present a few fun and insightful things about security incident response in the cloud.

Here’s what we decided. We focused on three challenges that we observed with organizations preparing for security incident response in the cloud, these are:

  • Skills: Cloud IR requires both solid security incident response skills and equally solid cloud native technologies skills
  • Joint nature: Many (but not all) cloud incidents will involve a CSP, 
    and many will involve a client, a cloud provider and one or more security service providers
  • Data: In many cases, telemetry, logs, traces data won’t be available or won’t be available via familiar mechanisms.

Next, we decided to focus on critical differences as well as similarities — no less critical — between security incident response on premise and in public cloud.

Now, if you want a one — line summary, the similarities mostly stem from the facts that the threat actors ultimately need to achieve their goals, and that responders need to know the environment to respond well (Duh, no brainer? Perhaps, but it affects how you do IR).

Here are the similarities:

  • Data preservation requirements.
  • Comprehensive understanding of the environment.
  • Standard investigative techniques have not changed.
  • Log data needs to be retained, normalized, and analyzed. Time Zones and time skew must be addressed.
  • Each incident is different.

Similarly, the differences mostly stem from the fact that cloud technology is often different, and operational practices for the teams behind called environments are different as well. As a side note, while people want to focus on logs from cloud services and containers, the fact that environments are just run differently, and jointly with your cloud provider partner, and that affects IR quite a lot.

Here are the differences:

  • Ephemeral & dynamic nature of the cloud.
  • Deep technical expertise of cloud native services required.
  • Different baseline and norms.
  • Log data retention, understanding, context and volume
  • Reliant on CSP & customer for relevant data.

Please watch the video and listen to the podcast. By the way, they cover completely different things, and in the podcast specifically, we share some deep secrets of how Google does IR in the cloud …

Related blogs on cloud security:


Originally published at Medium.

Thursday, November 17, 2022

More SRE Lessons for SOC: Simplicity Helps Security

As we discussed in our blogs, “Achieving Autonomic Security Operations: Reducing toil”, “Achieving Autonomic Security Operations…


As we discussed in our blogs, “Achieving Autonomic Security Operations: Reducing toil”, “Achieving Autonomic Security Operations: Automation as a Force Multiplier,” “Achieving Autonomic Security Operations: Why metrics matter (but not how you think)”, and the latest “More SRE Lessons for SOC: Release Engineering Ideas” your Security Operations Center (SOC) can learn a lot from what IT ops discovered during the Site Reliability Engineering (SRE) and DevOps revolution.

Let’s dive into another fascinating area of SRE wisdom that is deceptively simple — the principle of … simplicity (SRE book, Chapter 9 “Simplicity”). Say what? This sounds abstract and philosophical, how can it help my SOC today? Well, let’s find out!

The first point they make is a reminder of what makes it all exciting: “Software systems are inherently dynamic and unstable.” But SREs have it easy, here in our beloved realm of cyber, we don’t just have “systems that are inherently dynamic and unstable” but also attackers that affect them in a set of dynamic, beautiful (sometimes ugly) and unpredictable ways. 10X fun assured! So, yes, this is a big part of why security is fun, but also tricky. This means we need simplicity even more.

But what is simplicity? Phil’s 8 megatrends blog reminds us about this by calling one of his cloud megatrends “Simplicity: Cloud as an abstraction machine.” Specifically:

“A common concern about moving to the cloud is that it’s too complex. Admittedly, starting from scratch and learning all the features the cloud offers may seem daunting. Yet even today’s feature-rich cloud offerings are much simpler than prior on-prem environments — which are far less robust. […]
Cloud is only going to get simpler because the market rewards the cloud providers for abstraction and autonomic operations. In turn, this permits more scale and more use, creating a relentless hunt for abstraction. […]
The increased simplicity and abstraction permit more explicit assertion of security policy in more precise and expressive ways applied in the right context. Simply put, simplicity removes more potential surprise — and security issues are often rooted in surprise.”

Thus, removing surprises and reducing unique/broken/snowflake systems and silo’d processes will make security (and SOC in particular) easier. SREs have already figured much of this out.

To dive into the details, they say that their “job is to keep agility and stability in balance in the system.” For us in security, another (tricky) dimension gets added: the threat. Occasionally, compliance gets blended in too, and that may push the organization to old, ultimately more fragile ways of doing things (as a side note, compliance that reduces security by pushing for outdated approaches is a real thing). Fragile systems (and processes) + threats + regulations = a lot of complexity. And a complex system is always hard to monitor for threats, which are then also hard to investigate, making SOC life a pain.

Also, SREs say that “reliable processes tend to actually increase developer agility.” We wish we can always say that secure processes do the same. And, how about that, many actually do! Here at Google we have many examples of something that is secure and good for developers and good for business. Think well-implemented zero trust, that helps users, simplifies IT and reduces risk. It also makes the job of a SOC easier. Anyhow, we digress a bit…

Let’s look at more manifestations of the SRE principle of simplicity. Now, this is really juicy: “Essential complexity is the complexity inherent in a given situation that cannot be removed from a problem definition, whereas accidental complexity is more fluid and can be resolved with engineering effort.” This line alone is magical for the SOC!

I’ve always said, for example, that SIEM is complex, largely because its mission is complex. Now, if your SIEM vendor makes SIEM as complex as the mission, but not more complex, you may have a winner. Ideally, it should make it simpler, but frankly it won’t make it simple. Because it is simply not! However, excessive and removable complexity is a dire enemy of security.

Put another way, detection is hard, but some tools make it harder. Don’t use those, use the ones that don’t add complexity.Push back when accidental complexity is introduced” as SREs say. We definitely need to fight this battle in the SOC.

Further, if SREs say that “every new line of code written is a liability”, then in the SOC every detection rule is. You deploy this detection, and then you create work, toil in some cases, for you or your colleagues who have to respond to the resulting alerts. Think about it! The way to make this work, rather than fail, is to have a solid lifecycle for all detection content, in my view. Then every line you add delivers more value than liability, even if liability is never 0.

“The ability to make changes to parts of the system in isolation is essential to creating a supportable system.” So? To me, a bad “integrated” platform is worse than two good tools that can hook into each other via APIs. This is why I think best of breed ultimately won in security and suites and broad over-promising platforms lost (although this point is frankly very contentious, so let me quickly shuffle away from this particular argument…)

“Simplicity is an important goal for SREs, as it strongly correlates with reliability: simple software breaks less often and is easier and faster to fix when it does break. Simple systems are easier to understand, easier to maintain, and easier to test.” And of course simple systems and processes are easier to secure and monitor for threats. Now, some readers may say “but wait, I am a bank with 300 years of history, every process I have is complex, not simple.” Sure, but you still get to “push back when accidental complexity is introduced.” If your IT is inherently complex, the fight for reducing “excessive and removable” complexity is needed more, not less.

Naturally, simpler systems in your SOC help even more. Do you really need this rule with 5 correlated states or a playbook with 30 decision boxes? And if you have to make a 30 box alert triage process flowchart, then don’t make a 70 box flowchart?

To summarize, they say “software simplicity is a prerequisite to reliability.” We can add: also for security and threat “detectability” and “investigability” (can we just say observability?).

So, your mission, should you choose to accept it, is to push unneeded complexity out of your SOC. Where does complexity hide in your SOC? In detection content? Playbooks? Escalation processes? Workflows that involve other teams? Metrics and associated data collection? This is where you go and look at reducing the complexity.

Finally, “For SREs, simplicity is an end-to-end goal: it should extend beyond the code itself to the system architecture and the tools and processes used to manage the software lifecycle.“ I wish we had this for security and SOC in particular.

P.S. So I reread this post a few times (well, OK, more than a few times) and it still looks more conceptual than practical. So, perhaps one practical tip: when you encounter or create a SOC process, or a piece of technology in or around your SOC, think “does this add complexity?” and “is this complexity truly necessary?” If YES and NO, then think how to do things differently. If YES and YES, then think if the second question answer really is a YES…

Related blog posts:


Originally published at Medium.

Friday, November 11, 2022

Use Cloud Securely? What Does This Even Mean?!

An influential Gartner paper stated many years ago that “Clouds Are Secure: Are You Using Them Securely?”


An influential Gartner paper stated many years ago that “Clouds Are Secure: Are You Using Them Securely?”

So began the legend of cloud security vs secure clouds.

When I was an analyst, we sometimes had to discuss with clients whether various providers of public cloud services are “secure.” Over time, these discussions dwindled to a small trickle as clients ultimately saw enough evidence that cloud infrastructure is indeed radically more secure than most data centers. Admittedly, I still meet an occasional character who does not believe this, but I assure you, it is the truth. Moreover, cloud infrastructure is getting more secure all the time. However, security issues with client cloud environments did not dwindle to a corresponding trickle, and show few signs of such “dwindling.”

In fact, this situation led to the following truism: “Through 2025, 99% of cloud security failures will be the customer’s fault.” (source, also Gartner)

Thus, the explanation was always that the clouds are secure, but clients are not using them securely, and so they are to blame for the outcomes. As one of my current colleagues said on a call “this sounds a bit ‘blamy’” … and, yeah, it sure does.

So, what does USE CLOUD SECURELY mean in practical terms?

What entails secure use of the cloud?

What should one actually do to use the cloud securely?

What should one not do to use the cloud securely?

Naturally, I tried this first:

(source)

The results were interesting, but they mostly reinforced my impression that there is a wide confusion in the industry about what “use cloud securely” really means.

Next, I tried reading a whole load of materials such as “How to Make Cloud More Secure Than Your Own Data Center” , “Staying Secure in the Cloud Is a Shared Responsibility”, “Hybrid Cloud Security Best Practices” (and dozens of others, list available upon request, much of the content is paywalled, analysts need to eat).

This brought some clarity, but also an intense feeling of “it depends.” While not contesting this, I think answering this question with “do proper risk management for cloud environments” or “deploy appropriate security controls” is still not helpful for many. Some others said that this is just about the clients doing what is on their side of the shared responsibility model. However, we all know how well this works out in some cases.

A good (if a bit negative) frame emerged: “unless you are doing X, you are NOT using cloud securely” which for some values of X rings very true

A few islands of agreement emerged as well. These include:

  • You probably need to know what you have in the cloud (obviously, if you don’t, you are not using cloud securely)
  • You need a degree of awareness of real cloud threats and you need to know what threat model
  • You need to configure things securely (and actually know what this means), keep secure defaults, etc
  • You must get cloud IAM right for you (whatever that means, specifically), because ultimately you decide who can access your cloud, and not the provider.
  • You do need to detect threats against your cloud environment, your provider will help but ultimately you know your threat model better

Beyond the above, things start to look more fuzzy and agreement over what “use cloud securely” seems harder to reach.

So, conclusions:

  1. Before we scream at the clouds, we need more consensus and more certainty about what “use cloud securely” means in practical terms.
  2. Yet even after we achieve 1), there will be a lot of “it depends” (such as on risk appetite) left over; providers can help, but not magically do this (even though cloud providers can do more than what we do know, perhaps?)

Thoughts? Reactions?

Definitely expect more on this in the near future!

Related posts:


Originally published at Medium.

Monday, November 07, 2022

Anton’s Security Blog Quarterly Q4 2022

Great blog posts are sometimes hard to find (especially on Medium), so I decided to do a periodic list blog with my favorite posts of the…


Great blog posts are sometimes hard to find (especially on Medium), so I decided to do a periodic list blog with my favorite posts of the past quarter or so.

Here is the next one. The posts below are ranked by lifetime views. This covers both Anton on Security and my posts from Google Cloud blog, and our Cloud Security Podcast too (subscribe).

Top 5 most popular posts of all times (these ended up being the same as last quarter):

  1. “Security Correlation Then and Now: A Sad Truth About SIEM”
  2. “Can We Have “Detection as Code”?”
  3. “New Paper: “Future of the SOC: SOC People — Skills, Not Tiers”
  4. “New Paper: “Future of the SOC: Forces shaping modern security operations”
  5. “Beware: Clown-grade SOCs Still Abound”

Top 5 posts with the most Medium fans (these are also the same as last quarter):

  1. “Security Correlation Then and Now: A Sad Truth About SIEM”
  2. “Beware: Clown-grade SOCs Still Abound”
  3. “Can We Have “Detection as Code”?”
  4. “Why Is Threat Detection Hard?”
  5. “Stop Trying to Take Humans Out of SOC … Except … Wait… Wait… Wait…”

Top 5 Cloud Security Podcast by Google episodes:

  1. Episode 1“Confidentially Speaking”
  2. Episode 2 “Data Security in the Cloud”
  3. EP47 “Megatrends, Macro-changes, Microservices, Oh My! Changes in 2022 and Beyond in Cloud Security”
  4. Episode 3 Automate and/or Die?
  5. EP75 How We Scale Detection and Response at Google: Automation, Metrics, Toil

Random fun new posts:

  1. ”Why Your Security Data Lake Project Will … Well, Actually …”
  2. “Detection as Code? No, DETECTION AS COOKING”
  3. ”On Trust and Transparency in Detection”

Now, fun posts by topic.

Security operations / detection & response:

Data security:

Cloud security:

Enjoy!

Previous posts in this series:


Originally published at Medium.

Thursday, October 20, 2022

Why Your Security Data Lake Project Will … Well, Actually …

Long story why but I decided to revisit my 2018 blog titled “Why Your Security Data Lake Project Will FAIL!” That post was very fun to…


Long story why but I decided to revisit my 2018 blog titled “Why Your Security Data Lake Project Will FAIL!” That post was very fun to write and it continued to generate reactions over the years (like this one).

Just as I did when I revisited my 2015 SOC nuclear triad blog in 2020, I wanted to check if my opinions, views and positions from that time are still correct (spoiler: not exactly…)

As a reminder, the post stated that most organizations building DIY security data lakes would not succeed in these glamorous endeavors. I also predicted that years of immense (and costly) suffering would await the most persistent of them (and, yes, very, very few would succeed beautifully). In the end, most would either reach dramatically diminished goals or would spend the time/money and in fact accomplish nothing useful for security.

Note that this blog was informed by my observations of the previous wave of security data lakes (dating back to 2012) and related attempts by organizations to build security data science capabilities. While some think that this lakey excitement is recent, in reality, it dates back a decade or more.

So, in 2012, we said:

“Finally, “collect once — analyze many times for many goals” that include security, fraud and (rarely) operations model seems to be appearing in some leading organizations. Using the concept of a “data lake” where every team that needs the data can dip, apply schema at data read time (If needed) and solve a broad set of problems motivated them to explore big data approaches. Of course, the misguided motivation of “everybody else is doing it” (which is patently not true in this case) seems to be driving some Hadoop pilots and other dabbling with big data.“

However, we are not living in 2012 or 2018 anymore — we are in 2022. My attempts to start this discussion on Twitter revealed both proponents and opponents of the view that perhaps nothing has changed. So, has it?

Let’s review the arguments. I propose to use the following dimensions to decide on this: what changed and what stayed the same.

What I think is still true or unchanged:

  • Security people have not learned scaled data management and data analytics, namely building, maintaining, optimizing and evolving the data analytics stacks, and frankly using them effectively as well
  • Large scale data analysis infrastructure is still really hard, and actually still costly (well, this has caveats, see below)
  • Security data analytics talent shortage is still there, so if you have only a few people, they should use products, not build or maintain them (I used to joke around 2013 that the planet holds about 5 real security data scientists, two of whom are named Alex. Hi Alexes!)
  • Security (at least detection and response) is still a big data problem, and threat detection is still hard
  • Even a traditional SIEM (whether software or even SaaS) is hard for many organizations to operate and use (hence MDR and various managed models are still very popular)
  • Processes around security operations, and detection and response at many organizations are still very immature. Move to cloud have not changed this and sometimes set the clock back
  • Most threat detection still requires structured data and that means reliable collection, working parsers, data cleaning and other steps are still required, while key word searches only go so far.

What I think changed in significant ways:

  • Amazing new cloud data storage platforms are now available, and cloud data storage is not onerous to implement and run, compared to circa 2012 Hadoop
  • Many organizations have accepted cloud as the way to store large volumes of sensitive data; and some data more sensitive than logs
  • There is a lot more data we need to collect and analyze, such as from the public cloud environments
  • If well architected, cloud does allow for less costly scalable telemetry storage, coupled with dramatic decrease of costs to operate the system
  • There is much more sanity and less silly excitement about the role of ML in particular and data analytics in general for security. Machine learning works, but alien ML magic unicorns have not landed (sorry for mixing the metaphors here)
  • Cloud makes API integration between storage layer and security layer a lot easier and a lot more reliable; you no longer need to install two complex pieces of software and then pray for them to be friends. Data integration may be one API call away.
  • Thus, it is easier to create a workable and maintainable integration between the storage stack and security stack (I agree with Omer here, for sure, and I think this is the biggest position change in my thinking)

Still, to me, the decisive issue is whether decoupled SIEM is a good idea, and for what types of organizations. To remind, decoupled SIEM is where data storage technology (a security data lake, essentially) is built by one vendor while security brains are built by another.

For sure, today many people still prefer integrated, and use single vendor toolset or even an MDR (MDR itself, by the way, may well use a decoupled stack). For example, look at the 2022 SIEM Magic Quadrant, and integrated vendors still feature prominently. Still, my former colleagues have this to say:

“Data decentralization enables more cost-effective deployments, with more up-to-date data, and is expected to be a key trend for SIEM over the next 18 months”. (SIEM MQ 2022)

(my other sort-of colleagues agree and say this: “Top [SIEM] disruptor: security analytics on top of independent datastores.”)

So, back when my original data lake posts were written, integrated was winning and decoupled was losing — and losing badly at that. Today, integrated is doing fine, but I don’t think decoupled is losing anymore.

In light of this, I am ready to say that your security data lake project may still fail, but it probably does not have to ;-)

Thanks to my dear friend Augusto Barros for a very helpful discussion. Thanks to Omer Singer for indirectly motivating me to write this.


Originally published at Medium.

Monday, October 17, 2022

What is your Cloud SIEM Migration Approach?

This blog is written jointly with Konrads Klints.


This blog is written jointly with Konrads Klints.

TL;DR:

  • Migration from one SIEM to another raises the question of what to do with all the data in the old SIEM. A traditional approach was to let the old SIEM hardware languish until its data was no longer required.
  • When migrating from a cloud-based SIEM “A” to another cloud-based SIEM “B”, you have to contend that data is not easily transferable across SIEMs and/or that there will be significant data retention costs with the old SIEM (if you keep it up even without any new log collection)
  • A proposed solution to this would be to export data from SIEM “A” to a temporary data lake, and then use data analytics methods and serverless queries such as Google Dataproc or AWS Athena — for select use cases.
  • This requires only a moderate cloud and data analytics expertise which can be readily sourced if not available in-house.
  • Big cost savings come from keeping the data compressed in an inexpensive data storage system such as Google Storage nearline or similar.
  • The overall functionality is reduced as compared to a full blown SIEM, but still acceptable for common older data use cases — search by keywords for IR, IOCs during threat hunts or compliance data retrievals.

Problem statement

Deploying and using a Security Information and Event Management (SIEM) tool tends to occasionally generate negative emotions in people (no way, right!?). This sometimes leads to teams deciding to replace one product with another. As SIEM is simultaneously a security technology and a data management (and analytics) technology, one of the migration challenges is dealing with log data that has been accumulated, sometimes over a year or more.

Back in Anton’s analyst days, my advice to clients has been to NOT even try to migrate the log data. For example, this is what I said in a 2019 blog post:

There is no migration of collected log data, in most cases. It is just not worth the effort. Prepare to keep the old SIEM running for 9–12 months (some choose to do so unsupported, but YMMV) as a “legacy data store.” Now, you can try to export/convert/import, but this is messy, labor intensive, annoying and, frankly, can just be avoided by keeping the old data in your old SIEM or in some independent log repository (this reminds us why separate SIEM and CLM has been a good idea for many)”

I’d say this view originated in the age of software/appliance SIEM (say 1998–2010 or so) and has been accepted as conventional wisdom for SIEM migrations by most people. Organizations will keep an old database of SIEM data or a set of old SIEM appliances running, but not collecting data, just in case they need to see the data or in case some pesky auditor shows up and asks for it. Some older SIEM tools also allowed offline data archival, so one essentially needed a copy of the software and a bunch of tapes (yes, tapes!) that can — with some luck — be loaded back if needed.

Naturally, their support contract would have expired by then, and in some cases keeping the product running would be considered a license violation, but it was done regardless (after all, they were not using the product). Some would downgrade the license to the cheapest possible option and schedule decommission a year from now.

However, suppose you have decided to migrate from one SaaS or cloud-based SIEM platform to another (in this post, we used them interchangeably, but sometimes people do point out the differences). With a cloud-based SIEM you typically pay for data retention, thus potentially doubling your SIEM costs during the retention period — pay for new SIEM and pay for old SIEM (example)

For large deployments, this could be financially prohibitive. Also, this just sounds wrong, but what are the choices? Buy another, temporary SIEM? Use cheap (eh…not for petabytes) intermediate storage? Use open source (and then who pays for hardware or cloud storage)?

Or, perhaps we can pay somebody really clever to enable export / convert / import of the data? First, this requires your old SIEM to have the capability to bulk export. Second, SIEM tools don’t reliably support ingestion of “raw SIEM dumps” from another tool. Such raw SIEM dump usually looks like lines of text without metadata. SIEM engineers would have to write custom parsers for an unknown amount of log types with a high probability of errors in parsing. Even if the old SIEM can export raw logs, the process of ingestions can be really tricky, for a long list of reasons to painful to list here…

How can we inexpensively store security log files for long periods of time, while maintaining a reasonable, SIEM-like functionality?

Objectives

Now, what do we really want?

  • Have a searchable log storage for 1 year or longer (the number is driven by both IR needs as well as compliance), so we can retire the old SIEM immediately.
  • Able to support use cases such as simple hunt searches, incident response investigation and compliance data queries: search/retrieve logs by a few, common keywords such as IP address, domain name, username — essentially a substring match.
  • Must be able to select a date range for performance reasons (don’t search all logs if you want only to search logs in a specific date range, say from May to April 2021 )
  • Focus on select essential fields such as a hostname which would usually be present in the first 100 or bytes of any SIEM bulk log dump lines and be somewhat consistent across logs types

Couldn’t we just OpenSearch/ELK it?

One approach would be to somehow export the data from the old SIEM and then shove all of the data into an ElasticSearch/OpenSearch instance with minimal pre-processing. At first glance, there are several advantages to this: ELK is a well understood solution that security engineering teams could reasonably cope with. A small, 3–4 node cluster with ample storage seemingly would fit the bill.

The cost breakdown is as follows:

  • Storage: the “compression” ratio of raw vs on-disk in Elastic is best case 50% so, a 10TB raw dump would require at least 5TB hot storage attached to a VM which would cost $2000 per month or $24,000 per year. More likely the storage requirements will be higher.
  • Compute: 4 instances of e2-highmem-4 at $131 per month per instance or $6228 per year (with ~40% discount for committed use it will come down to less than $4000 per year).

We don’t expect the performance to be amazing (because ELK), but for the occasional search this should be just fine. The above example may set you back mid-five digits (say up to $50K) plus the upfront cost of engineering time which will vary depending on how many fields you will want to extract (this could be as low as one). That’s not overall terrible. However, if you have petabytes of data, it really is that. Terrible!

The main advantage of this approach is that it feels familiar, there’s the search box and the time slider.

Using Google Cloud Platform — A Cheaper, Serverless Approach

It’s 2022 and we live in a world with flexible computing capabilities. Surely we could engineer something that eats less money while offering a lot of flexibility and some analyst creature comforts — an experience that is better than grep?

Google Dataproc serverless with Google Cloud storage backend seemingly meets most of these needs. It has a low running cost and Spark Notebooks allow analysts to make use of SQL-like queries.

Experience shows that unless the goal is to extract all fields from logs (not what we are aiming at here), then a regex’y solution for the desired types of events is usually good enough. This requires moderate data analytics and cloud expertise, perhaps someone with ~3–5 years of experience in the field.

Economics

  • Google Cloud Platform Hot Storage is $0.2 per GB/month or $2500 per TB/year. Nearline storage is 50% cheaper.
  • Conservative log compression ratio of 10:1 using gzip. So 1TB of actual storage would be only $2500/y.
  • You only pay for compute capacity when you use it using Dataproc serverless.
  • If you keep data in hot storage and do not move out of the same region GCP service, then data transfer is free.

An AWS Athena approach

Similarly, AWS Athena is a serverless bulk data query method. The user is charged for storage and only for the query execution time. Under the hood, Athena is a wrapper around the Presto engine. This makes it very attractive to store data that isn’t used all the time — just our case here. We followed a naive approach of no pre-processing of the data and querying it using the runtime extraction — regular expressions. We uploaded data to S3. From the user’s perspective as long as the wait time isn’t too big, there are no incentives to optimize query performance.

Economics

  • You are billed only for data transfer from S3. This means we don’t care to make queries optimized.
  • No costs associated with computing. Cool!

We encountered several problems along the way:

  • Athena doesn’t understand ISO8859 formatted timestamps nor it understands timezones. This means that data either needs to be preprocessed to bring it to a common, agreed time zone such as UTC or queries must take it into account. Not a major obstacle, but it highlights that some data analytics skills are required for the whole export-SIEM-to-serverless approach.
  • One means of cost reduction per query is to store data in a partitioned fashion in S3 in a hierarchical folder structure: there would be an increased performance and a reduction of cost as Athena wouldn’t have to scan the entire data set. However, the data would have to be arranged in the hierarchical format before which may require some pre-processing.

Overall, this approach also works.

Conclusions

Migration of one SIEM to another raises the question of what to do with data in the old SIEM. A traditional approach was to let the old SIEM and hardware languish until its data was no longer required.

A solution to this would be to export data from SIEM “A” as a log dump and use data analytics methods — data lake and serverless queries such as Google Dataproc serverless or AWS Athena — for select use cases. This requires only a moderate cloud and data analytics expertise which can be readily sourced if not available in-house.

The overall functionality is reduced as compared to a full blown SIEM, but still acceptable for typical older data use cases — search by keywords for IR, IOCs during threat hunts or compliance data retrievals.

Finally, one can use the same approach for migration from on-prem SIEM, provided the data can be uploaded to the cloud cost-effectively.

P.S. All cost estimates are essentially somewhat educated guesses; do your own math on your own data, please.


Originally published at Medium.

Thursday, October 13, 2022

This is awesome, let me see if I can create a gentlemanly "semi-counter-argument: and post it in a…


This is awesome, let me see if I can create a gentlemanly "semi-counter-argument: and post it in a few days.


Originally published at Medium.

Wednesday, October 12, 2022

Google Cybersecurity Action Team Threat Horizons Report #4 Is Out!

This is my completely informal, uncertified, unreviewed and otherwise completely unofficial blog inspired by my reading of our fourth…


This is my completely informal, uncertified, unreviewed and otherwise completely unofficial blog inspired by my reading of our fourth Threat Horizons Report (full version) that we just released (the official blog for #1 report, my unofficial blog for #2, my unofficial blog for #3).

My favorite quotes from the report follow below:

  • “in Q2 threat actors frequently targeted weak and default-password issues for initial compromise, factoring in over half of identified Incidents.” [A.C. — that is not ‘Q2 of 1998’, that is 2022! Especially the ‘no credentials’ part, which to me smells like the 1980s, not even the 1990s]
Cloud Compromise Factors from TH #4 — Google Cloud
  • “Once inside, threat actors frequently engaged in cryptomining, accounting for nearly two-thirds of incidents (65%).” and “cryptominer attacks are often partially or fully automated, dramatically reducing their time to exploit an available vulnerability.“
  • “The high level of SSH activity suggests that organizations are using either no credentials or default credentials when spinning up cloud instances.” [A.C. — somebody from the IR firm whose name starts with ‘M’ told me the other week that ‘in the cloud, they still have the mid-2000s in regards to some security practices’ and this is a great, if sad, example of that!]
  • “The controls fail to identify the malware assets’ nefarious nature as they check the assets’ context and external characteristics, instead of exploring their content in more depth.” and “10% of well-known, popular websites are seen to be distributing malware. Malware “legitimacy” is inherited from credible hosts.“ [A.C. — this to me is a fun reminder that naïve badness ‘blocklists’ fail]
  • There is a lot of signed malware because “attackers often fraudulently accessed signing workflows or signing authorities to sign their code — increasing the likelihood of its downstream acceptance” [A.C. — the news here is not that they do, but that nasty “o” word — “often”]
  • “Kimsuky, a nation-state threat actor, has been observed by researchers at Volexity accessing user Gmail account data through a hidden Chrome browser extension known as SHARPEXT. The group […] was able to install a malicious browser extension via phishing, leveraging pre-authenticated browser activity to read and exfiltrate data from other services such as Gmail content” and “installation of a developer-mode browser extension which, through a DevTools workaround, has its security warnings suppressed and targets a user’s cloud-accessed data” [A.C. — this is reasonably notable, and fairly scary too!]
  • … and a useful reminder here that ”increased productivity provided by seamless SSO also provides broader access for attackers to otherwise confidential data.” [A.C. — ‘login once — get everywhere’ [if done wrong] heps both the good and bad actors, unless zero trust is also done well]
  • “they brute-forced the instance’s password and enrolled their own device in the NGO’s multi-factor authentication (MFA) process[A.C. — another reminder that MFA is not useful if anybody can enroll a malicious device into it]
  • “Threat groups have been observed leveraging compromised service account credentials to run expensive cryptomining workloads in customer environments, but greater concern would arise should they choose to keep these actions covert and leverage the access for other nefarious activities. ” [A.C. — to me, this reminds us that relying on attackers being very noisy for detection is not a great strategy]

Now, go and read the report!

Related posts:


Originally published at Medium.

Dr Anton Chuvakin