Wednesday, March 16, 2022

How to SLO Your SOC Right? More SRE Wisdom for Your SOC!

As we discussed in “Achieving Autonomic Security Operations: Reducing toil” (or it’s early version “Kill SOC Toil, Do SOC Eng”) and…


As we discussed in “Achieving Autonomic Security Operations: Reducing toil” (or it’s early version “Kill SOC Toil, Do SOC Eng”) and “Stealing More SRE Ideas for Your SOC”, your Security Operations Center (SOC) can learn a lot from what IT operations learned during the SRE revolution. In this post of the series, we plan to extract the lessons for your SOC centered on another SRE principle — Service Level Objectives (SLOs).

In brief, this is about metrics. SOC metrics have long fascinated me, and this is a chance to learn from a new domain that is generally ahead of security in its systems thinking.

Before we go there, what’s an SLO? “An SLO is a service level objective: a target value or range of values for a service level that is measured by an SLI. “ OK, what’s an SLI? Well, “An SLI is a service level indicator — a carefully defined quantitative measure of some aspect of the level of service that is provided.” (all quotes are from the SRE book here)

So? We measure something (SLI) and we set the target value (SLO). Now, what about people who have only heard of SLAs? Well, SLA is an agreement about the above: “an explicit or implicit contract with your users that includes consequences of meeting (or missing) the SLOs they contain.”

I am not going to spout clichés like “what gets measured gets done” here, but metrics and SLIs/SLOs will to a large extent determine the fate of your SOC. My favorite, if a bit dated, example is: SOCs (including at some MSSPs) that obsessively focus on “time to address the alert” (that they naively consider to be the same as MTTD, BTW … WTF) end up radically reducing their security effectiveness while making things go “whoosh” fast. If you equate MTTD with “time to address the alert” and then push the analyst to shorten this time, you will not have a good time … while the attacker will.

So, yes, SREs also start the SLO discussion with the reminder that “choosing appropriate metrics helps to drive the right action.”

Now, a naïve view of metrics would be that “whatever sounds bad” (problems per second, incidents per employee, etc) need to be minimized while “whatever sounds good” (successes, reliability, uptime, etc) need to be maximized … ad infinitum. But hey… here is a new insight: sometimes good metrics have an optimum level, and yes, even reliability (and maybe even security). Read the SLO chapter in the book for a full example, but they have an example of a service where the reliability was too high. How is it bad? “Its high reliability provided a false sense of security because the services could not function appropriately when the service was unavailable, however rarely that occurred. […] SRE makes sure that global service meets, but does not significantly exceed, its service level objective.”

There is a fun SOC lesson here: some security metrics have optimum value. The above-mentioned time to detect, I bet, has an optimum for your organization at least, if not perhaps a global optimum (similar to my patch sound barrier). Another example: the number of phishing incidents — that screams to be pushed to 0 , right? — may have an optimum too: if nobody phishes you, this is probably because they already have credentialed access to many of your systems. So in your SOC, think of SLI optimums, and don’t automatically assume 0 or infinite for metrics.

The SRE book reminds us that “good metrics” may need to be balanced with other metrics, rather than blindly pushed up. “User-facing serving systems generally care about availability, latency, and throughput. […] Storage systems often emphasize latency, availability, and durability. […] Big data systems, such as data processing pipelines, tend to care about throughput and end-to-end latency.” In a SOC, this may mean that you can detect fast, review all context, perform deep threat research — but the balance may differ for various threats and situations. So, think combinations of metrics, not mere numbers.

Another lesson from SREs of value to your SOC: “Whether or not a particular service has an SLA, it’s valuable to define SLIs and SLOs and use them to manage the service.“ Indeed, I agree that SLIs and SLOs matter more for your SOC then any agreements i.e SLAs. Metrics and targets before handshakes!

Now, here is a gem: “Most metrics are better thought of as distributions rather than averages.” For you, my sole statistically skilled reader, this is obvious. For others: what do you make of an average alert response of 20 minutes? Is this “all alerts are addressed in 18–22 minutes” or “all alerts are addressed in 5 minutes, while 1 alert is addressed in 6 hours”?

In fact, “The higher the variance in response times, the more the typical user experience is affected by long-tail behavior.” This is definitely something I’ve seen in SOCs — that one outlier event is probably the one that matters most. To this, SRE advice is “Using percentiles for indicators allows you to consider the shape of the distribution.”

The other epically useful concept from SRE is of course “the error budget.” This foundational concept may not be clear to my security peers so here is SRE advice verbatim: “allow an error budget — a rate at which the SLOs can be missed — and track that on a daily or weekly basis. (An error budget is just an SLO for meeting other SLOs!)

SOC value here is not immediately obvious, but here it is: this is security, and the game is not about the metrics, ultimately, it is about the threat actor. I’d rather miss the SLO, but get the threat in my environment. I’d rather spend more time than comply with rigid time metrics. Ultimately, the defenders win when the attacker loses, not when the defenders “comply with an SLA.” Error budget concept is your friend here.

Further, the SRE thinking goes like this: “It’s both unrealistic and undesirable to insist that SLOs will be met 100% of the time: doing so can reduce the rate of innovation and deployment.“ More broadly, and as we say in our recent paper with Deloitte on SOC (“Future Of The SOC: Process Consistency and Creativity: a Delicate Balance”), “this adherence to process and lack of ability for the SOC to think critically and creativity provides potential attackers with another opportunity to successfully exploit a vulnerability within the environment, no matter how well planned the supporting processes are.“ (read more here)

Now, here is another brain teaser from our SRE brethren: “Don’t pick a target based on current performance.” Huh? This is very common, in my experience, so is this really that bad? Let’s see, an analyst handles 30 alerts a day (SLI), their manager wants to improve by 15% so they set the SLO to 35 alerts a day. All good? Well, wait a second. How many alerts are there? Leaving aside the question of whether it is the right SLI for your SOC (spoiler: it is not), what if you have 5000 alerts, and you drop 4970 of them on the floor. When you “improve” you will drop merely 4965 on the floor. Is this a good SLO? No, you need to hire, automate, filter, tune or change other things in your SOC, not set better SLO targets.

To this, our SRE peers say: “As a result, we’ve sometimes found that working from desired objectives backward to specific indicators works better than choosing indicators and then coming up with targets.“ AND “Start by thinking about (or finding out!) what your users care about, not what you can measure.” In the SOC, this probably means start with threat models and use cases, not the current alert pipeline performance.

Now, here is a cryptic one: how many metrics do I need in my SOC? SREs wax philosophical here: “Choose just enough SLOs to provide good coverage of your system’s attributes.” In my experience, I’ve not seen people succeed with more than 10, and I’ve not seen people describe and optimize SOC performance with less than 3. In other words, I don’t know. However, SREs offer a neat test: “if you can’t ever win a conversation about priorities by quoting a particular SLO, it’s probably not worth having that SLO.”

SLOs will get to define your SOC so define them the way you want your SOC to be: It’s better to start with a loose target that you tighten than to choose an overly strict target that has to be relaxed when you discover it’s unattainable. SLOs can — and should — be a major driver in prioritizing work for SREs and product developers, because they reflect what users care about.” This of course applies verbatim in a SOC.

Finally, make SLOs for your SOC public (well, within the company) just as SREs say (Publishing SLOs sets expectations for system behavior.) The benefit is that nobody can blame you for non-performance if you perform to those agreed upon SLOs (that may became SLAs).

At the end, if you are tired of reading my ramblings, here is a SOC metrics resource list I solicited from the community:

(for other advice and specific metrics, mine this thread)

Related posts:


Originally published at Medium.

Thursday, March 03, 2022

Anton’s Security Blog Quarterly Q1 2022

Great old blog posts are sometimes hard to find (especially on Medium) , so I decided to do a periodic list blog with my favorite posts of…


Great old blog posts are sometimes hard to find (especially on Medium) , so I decided to do a periodic list blog with my favorite posts of the past quarter or so.

Here is the next one. The posts below are ranked by lifetime views. This covers both Anton on Security and my posts from Google Cloud blog, and our Cloud Security Podcast too (subscribe).

Top 5 most popular posts of all times:

  1. “Security Correlation Then and Now: A Sad Truth About SIEM”
  2. “Can We Have “Detection as Code”?”
  3. “New Paper: “Future of the SOC: SOC People — Skills, Not Tiers”
  4. “Beware: Clown-grade SOCs Still Abound””
  5. “New Paper: “Future of the SOC: Forces shaping modern security operations”

Top 5 posts with the most Medium fans:

  1. “Security Correlation Then and Now: A Sad Truth About SIEM”
  2. “Beware: Clown-grade SOCs Still Abound”
  3. “Can We Have “Detection as Code”?”
  4. “Why Is Threat Detection Hard?”
  5. “A SOC Tried To Detect Threats in the Cloud … You Won’t Believe What Happened Next”

Top 5 Cloud Security Podcast by Google episodes:

  1. Episode 1“Confidentially Speaking”
  2. Episode 2 “Data Security in the Cloud”
  3. Episode 8 “Zero Trust: Fast Forward from 2010 to 2021”
  4. Episode 27 “The Mysteries of Detection Engineering: Revealed!”
  5. Episode 17 “Modern Threat Detection at Google”

Random fun new posts:

  1. “Anton and The Great XDR Debate, Part 3”
  2. “Left of SIEM? Right of SIEM? Get It Right!”
  3. “Kill SOC Toil, Do SOC Eng”

Now, top posts by topic.

Security operations / detection & response:

Data security:

Cloud security:

Enjoy!

Previous posts in this series:


Originally published at Medium.

Tuesday, February 22, 2022

Anton and The Great XDR Debate, Part 3

TLDR: no, this post still does not contain the Ultimate Answer for XDR, Life and Everything Question. Moreover, I don’t think anything ever…


TLDR: no, this post still does not contain the Ultimate Answer for XDR, Life and Everything Question. Moreover, I don’t think anything ever will. While we discuss XDR, the market forces change the definitions, vendors pivot away, analysts ponder, customers cry… well, the cyber-usual.

To start, I’ve had many conversations about XDR recently. Some were the ones where I sought answers, while others were where I sought questions and some were where people sought answers from me.

Now, I am well aware that this debate does not really touch the needs of many real security practitioners. Somebody asked me on social media why I am so obsessed with XDR. To me, the answer is I need clarity in technologies that we deploy. The clarity is essential to match products to requirements, to compare tools, and to cover the gaps in detection and response posture (and in security in general).

As you remember from my excellent Part 1 and from my — yeah, I know — mediocre Part 2, XDR remains a mystery to a whole lot of people. So, philosophically, I don’t want things to be confusing in an area where people are supposed to spend real money and to reduce real risks to their organizations.

So in recent days, my journey to XDR clarity has led me back to SIEM, SOAR and EDR. Specifically, one vision of XDR is that of consolidation married to simplification. Or, as I said in one private conversation, XDR as an integrated platform of minimized components.

First, a humorous take on this:

(source)

Now, XDR is NOT SIEM + SOAR + EDR. That would just be mad. However…. XDR may in fact be about

“SIEM -”

+

“SOAR -”

+

“EDR -”

What do the minus signs stand for? In my mind, they stand for reduced complexity, narrower (more focused) functionality and minimized frictions.

This vision of XDR seems more sane to me than “XDR as an improved EDR.” The slogan of “consolidate while slicing complexity” will probably have a lot more fans then “extend the endpoint technology to, well, not endpoint” 

This view of XDR is not my invention, even though the framing probably is. Reading what my former colleagues wrote recently (and this too), for example: “Extended detection and response is a platform that integrates, correlates and contextualizes data and alerts from multiple security prevention, detection and response components. XDR is a cloud-delivered technology comprising multiple point solutions and advanced analytics to correlate alerts from multiple sources into incidents from weaker individual signals to create more accurate detections. “ and “Use use-case analysis to improve security operations center (SOC) productivity and accuracy, or for risk reduction to help justify the addition of an XDR solution.”

In the above, they didn’t explicitly call the simplification or minimization of components, but they do mention narrower mission (e.g. SIEM is supposed to handle threat detection and compliance, while XDR has nothing to do with rules and regulations or insider threats for that matter).

To remind, the word “integrated” has a bad history in our industry. So far, everybody who promised an integrated security platform essentially failed or was found to be a bad liar. This has been the case since the 1990s, as I recall. Now people may want a more integrated experience (such as around their SIEM), but promising “all in ones” and “single pane of glasses” generally has an abysmal track record [well, the promising was fine, it is the delivery that was problematic :-)]

So, XDR is NOT an integrated platform of the stuff you already have. However, that XDR may be an integrated platform of several key pieces that were simplified, minimized, focused and then integrated. So, X may mean “eXcised”, not “eXtended” or “eXpanded”…

Now, the details are up to the vendors, but a log manager or a simple SIEM married to some endpoint visibility and canned detections coupled with workable response playbooks may be a valuable bundle. Simple and focused on a narrow mission and without any scope creep! We can even call it XDR and I can see people willing to buy that….

What do you think?

So, frankly, I don’t know what XDR is today. I know many people who think they do — and most of them don’t agree with each other. Review the technology presented to you and match it to your use cases and threats, don’t obsess about the buzzwords. Get a good cloud SIEM :-)

End of the story?

Related:


Originally published at Medium.

Tuesday, February 15, 2022

Google Cybersecurity Action Team Threat Horizons Report #2 Is Out!

This is my completely informal, uncertified, unreviewed and otherwise unofficial blog inspired by my reading of our second Threat Horizons…


This is my completely informal, uncertified, unreviewed and otherwise unofficial blog inspired by my reading of our second Threat Horizons Report (full version, short version) that we just released (the official blog for #1 is here).

Google Cybersecurity Action Team

My favorite quotes follow below:

  • “Threat actors have been known to use tools native to the Cloud environment rather than downloading custom malware or scripts to avoid detection. This “living off the land” technique has been used extensively in on-premise compromises and is being adopted in Cloud environments. ” [A.C. — highlight is mine, just a reminder that “malware-less” intrusions are the norm in the cloud too]
  • Sliver is an open-source, cross-platform adversary emulation framework that allows adversaries to deploy and control implants on victims’ Windows, Mac, or Linux computers from a central coordinating server.” [A.C. — not all badness is about Cobalt Strike …]
  • “Over the past 12+ months, the actor has launched multiple campaigns against the security and vulnerability research community including the following techniques: Developing fake social media profiles and submitting real bugs to bug bounty programs in order to build credibility. […] Suspected of using 0-days, which were stolen from some of their victims. ” [A.C. — imagine being hit by a 0-day that the attacker stole somewhere …]
  • “Google Cloud is continuing to see scanning (400K times a day) and expects similar, if not more scanning levels against all providers, and so we recommend continued vigilance in ensuring patching is effective. ” [A.C. — so, no, log4j is not over]
  • “The availability of outbound connections to conduct a reverse SSH tunneling from the Cloud Shell to any endpoint on the Internet is serving as a means for threat actors to distribute malicious campaigns or perform harmful activity.” [A.C. — this is a reminder that an outbound connection is always an ominous sign, cloud or not]

Enjoy the report!

Related:


Originally published at Medium.

Tuesday, February 08, 2022

Who Does What In Cloud Threat Detection?

This post is a somewhat random exploration of the cloud shared responsibility model relationship to cloud threat detection.


This post is a somewhat random exploration of the cloud shared responsibility model relationship to cloud threat detection.

Funny enough, some popular shared responsibility model visuals don’t even include detection, response or security operations. Mildly embarrassing, that.

Anyhow, let’s start here: a naïve view of shared responsibility model and detection is simply the following: the cloud provider (CSP) is responsible for detecting threats to their backend systems while the customer is responsible for detecting threats to whatever they put in the cloud. This view is often promoted by people who have never seen a cloud and who probably cannot even spell “I-D-S” :-)

So, if you think about it for 5 seconds, you realize that this is naïve and frankly misguided. Or, at least, simplistic bordering on impractical. In the real world, very clearly there is a role for cloud providers in both facilitating threat detection against whatever the customer puts in the cloud (say via a network IDS), or in some cases, just doing that detection themselves for customers (like say what we do here or here).

Thus, the lines for shared responsibility for detection are drawn somewhere else. We need to go look for those lines and/or redraw them if they are not visible anymore. By the way, we need to go look for them, because many cloud security issues originate at the seams of the shared responsibility model or arise due to cloud customers assuming too much about what the CSP is responsible for …

Let’s look at some examples first, because I feel this is a good place to do “bottom up” thinking.

A CSP develops and offers tools for detection, while customers are responsible for configuring the tools, including deploying the correct rules and detection content; then customers handle the alerts. This feels normal and largely parallel to what we see on-premise, no? However, the attack surface of a modern public cloud is so fundamentally different from the attack surface of an on-premise data center that translating from one to the other is hard. Moreover, many organizations using cloud are short-staffed (to put it mildly) in their security departments and may not have anybody who knows what to look for and what detection logic to build for the cloud. Frankly, even decent SOCs are often confused about cloud detection. So just as we rightly impugn “lift and shift” of workloads, we should neither lift and shift our threat models nor our detection rules. Of course, you know what this means: more work …

A CSP develops and offers tools that come complete with detection logic to detect common threats against the customer side of the cloud; the customer still gets the alerts. This approach of course makes the detection work like magic (for a customer), but limits the ability to tune and adjust the detection use cases based on your business in the cloud. To me, this means that fully shifting the detection responsibility to a CSP is a bad idea today (so the responsibility line is not drawn such that CSP is doing all the work). Note that here and in the previous scenario, the managed service provider (MSSP / MDR) may get the alerts to help the client figure out what to do (if they know how to handle the cloud themselves, that is).

A CSP develops infrastructure for detection which is then used by a third-party vendor who then sells detection tools complete with content / rules to the cloud users. This adds yet another party in the shared responsibility model. Alerts are handled by a client or by some managed services somewhere. In this case, we do have to draw more lines: CSP responsibility line, client responsibility line, 3rd party technology vendor responsibility line and potentially the managed service responsibility line (a lot of line drawing afoot!). While the third party vendor may do a better job than a CSP in some areas, they won’t understand the underlying technology as well (it also further divorces system creation from rule content creation that may be a recipe for unhappiness at times).

Naturally, a CSP also develops and operates the detection tools that detect threats to their infrastructure (and handle these particular alerts); here the naïve view is essentially correct, at least in part.

So, what do we learn here? Let me try to “overthink” it a little bit and present a table that summarizes the routes we just discussed.

Anton’s Cloud Threat Detection Table

(BTW, secretly, I think that in the long run the CSPs will do more of this work; this is why I am not investing any money into multi-$B cloud security startups).

Finally, some of you have skimmed this post and now have a burning question: which route is “better”? Furthermore, some of you are going one step further: which route is better for what cloud models, services, migration approaches? These are great questions and they will make a great follow-up post; this was meant to frame and start, not provide the final answer …

Thanks to Tim Peacock for helpful comments and some text contributions. Thanks to Anna Belak for the initial inspiration for this post.

Related blogs:


Originally published at Medium.

Friday, January 21, 2022

"Send me all your log data, and I'll tell you what's happening" (1996-) ?


"Send me all your log data, and I'll tell you what's happening" (1996-) ? 2026? 2036? Something else...


Originally published at Medium.

Yes, and this pre-dates the ML craze by a decade and a half...


Yes, and this pre-dates the ML craze by a decade and a half...


Originally published at Medium.

"we believed (wrongly) that by centralizing data we could (maybe by magic) analyze it faster" <…


"we believed (wrongly) that by centralizing data we could (maybe by magic) analyze it faster" <- indeed. We did believe that centralization is a BIG step to this, and it was a step. A small one :-)


Originally published at Medium.

Symantec SIM (well, acquired) was a sad piece of tech with a few cool features, frankly.


Symantec SIM (well, acquired) was a sad piece of tech with a few cool features, frankly.


Originally published at Medium.

Thursday, January 20, 2022

20 Years of SIEM: Celebrating My Dubious Anniversary

On Jan 20, 2002, exactly 20 years ago, I joined a “SIM” vendor that shall remain nameless, but is easy to figure out. That windy winter day…


20 years of SIEM?

On Jan 20, 2002, exactly 20 years ago, I joined a “SIM” vendor that shall remain nameless, but is easy to figure out. That windy winter day in northern New Jersey definitely set my security career on a new course.

With this post, I wanted to briefly reflect on this ominous anniversary. Where do we begin?

Let’s start with a sad fact that some of the problems that plagued the SIM/SEM of late 1990s and early 2000s are still with us today in 2022. One of the most notorious and painful problems that has amazing staying power is of course that of data collection.

I remember how our engineers struggled in 2002 with some API-based collection from a known firewall vendor. API-based log collection seemed new and weird back in the day (“why can’t they do UDP 514 syslog like normal people?”). Today, the current generation of engineers still struggles with some cloud-based collection mechanisms for telemetry data … and that is even before observability for security truly arrives.

Another problem — as we know now — possesses amazing staying power. “We collect data and then you get the insights” promise was made by SIEM vendors as early as 1999.

Do the SIEM operators today feel they got what was promised? Generally, no. In some areas, we are worse off, as users need to work more, not less, on getting the value from their collected data. Personally, I still hear SIEM users complaining loudly about “it just collects data” and “it is hard to get the outcomes we want.” This happens after a quarter of a century of promising that! Note to the excited security data lakers here: sorry, guys, but you are making this one worse, not better …

Next, another reminder is useful: the birth of security information management and security event management definitely predates the rise of compliance. In recent years, it always annoyed me when people equated SIEM with compliance, because I lived through the era before compliance mandates became “cool” in security. Thinking back to 2002, SOX just came out, HIPAA was new and cool, while PCI DSS … was not born yet.

If you are curious, what did people care about those days? Here are some juicy quotes from SIM / SEM / NSM / ITSM / LEM marketing of that era. BTW, lots of names for this technology space were in use back then, ultimately SIM/SEM won and then was forever fused together by Gartner.

(date: 2002, source)

(date: 2002, source)

(date: 2002, source)

(date: 2002, source)

(date: 2002, source)

(data: 2003, source)

[BTW, this last historical artifact makes me a bit angry, because at least one of the “cool” “search-based ‘SIEM’” vendors cannot do this today, in 2022, while the cheapest and simplest of the 1st generation SIM/SEMs could do it in 2003. This is literally 20 years of regress in front of your eyes!]

My writing at the time shows that I was quite obsessed with correlation, especially correlation of normalized / categorized data that allowed detection content (rule) creation without knowing the event IDs and other details from each log source [because single event matching was not considered cool in 2003, and so why is it acceptable in 2022?]

I have not kept copies of my oldest presentations (very few like this gem from 2003 survived), but my old speaker page reminds me that I focused on “Log Analysis for Security” (2004), “What Every Organization Should Monitor and Log” (2005) and “Log Mining for Security” (2006).

The point I’m making with these fond memories is that a lot of the problems we struggled with in our security operations are the same problems we struggle with today. Definitely, the technology context has changed, but the security challenges remain the same to a very large extent.

As an example, alert overload was what gave birth to SIM (usually network IDS alert overload back then). Guess what problem we are discussing now 20 years later? Alert overload of course!

One of the things that, in my estimation, the early SIEM did better than the later, text search-based products was focus on the quality of data. There were a lot of naïve decisions made in the early days (some for good reasons, some for bad, some because of technology limitations) such as to focus on data that is deemed security-relevant and discarding the rest. I recall agonizing over some Cisco event IDs when I was working with our log source integration lab.

Now, the more cynical among you may remind me that early SIEMs were plagued by parsing errors (yes, they were!) and that data quality was — in that sense — inferior to that of a tool that collects raw logs. Sadly, this is correct (this made it hard to trust those SIEMs), so perhaps this is a debate for another time. One approach I tried to use to solve it was to advocate the case for log standards.

XDR fans today may find it distasteful, but early SIEM pioneers really wanted to deliver effective and usable detection rules out of the box. Back in the day, we almost always called them correlation rules, because we that to be a good product in the space you have to have stateful correlation and not just simple matching.

As another theme, early on I was right in stating that workflow matters a lot for a good SIEM product. This happened perhaps 10 years before the first SOAR product was born. I recall designing a workflow system for our SIEM assuming people in the SOC will live in that UI most of the time…

In the early days, I also spent a fair bit of time looking at how threat actors behaved and then writing detection content (I even ran a honeypot and then used the observation to create correlation rules). Naturally, that predated the age of APT so most of our detection content was focused on commodity actors — script kiddies as they were known back then. But hey — it wasn’t the auditors!

After a few years (2006-ish), I spotted that a complete collection of logs would become a thing and left my original SIEM employer. Note that I did not expect that collecting raw logs would kill the original normalized, enriched and categorized SIEM… I wanted to have both full/raw and clean/enriched data (kinda what we have now…)

After that, I stepped away from SIEM and had a chance to experience a very well built SaaS security tool — although not in the SIEM area. This made me excited about the chance of creating a cloud-native or SaaS SIEM. Of course, now everybody knows this. Back in 2009, SaaS SIEM didn’t really exist…but that is the story for another day.

Now, I am not sure I can be called an official SIEM historian (I may not be the oldest SIEM developer alive, but it is possible I am the one who tolerated SIEM in my life for the longest time...), but here is what I have.

Feel free to share your SIEM / SIM /SEM or log analysis stories, let’s have a party! :-)


Originally published at Medium.

Wednesday, January 12, 2022

Left of SIEM? Right of SIEM? Get It Right!

This post is perhaps a little basic for true SIEM literati, but it covers an interesting idea about SIEM’s role in today’s security. I…


This post is perhaps a little basic for true SIEM literati, but it covers an interesting idea about SIEM’s role in today’s security. I suspect that this topic will become even more fascinating in light of the appearance of XDR — but more on this a bit later…

So let’s talk about what’s to the left and to the right of SIEM. Note that this has nothing to do with the “shift left” of software development. Well, it has nothing to do with it directly, but perhaps there are lessons to be learned in this area for our domain.

Your SIEM looks at certain inputs and produces some outputs (this is not all about GIGO, BTW). But, yes, if the inputs are dirty and/or the outputs drop on the floor, SIEM cannot succeed, no matter what the SIEM vendor says or does.

When I say that somebody succeeded with SIEM, it often implies that they got things to the left of SIEM and to the right of SIEM right or at least “right enough.” It is not enough - absolutely, not enough! — to just install your SIEM software correctly or sign up for a cloud SIEM service.

Ultimately, I believe that success with SIEM means succeeding with what’s left and what’s right. Because your goal is not to “succeed with SIEM” at all — but to succeed with detection and response, compliance monitoring or whatever you use the tool for.

By the way, why do I offer this type of structured thinking about what is to the left and to the right of your SIEM? In my opinion, this approach will help make your SIEM operation more effective and will help you avoid some still-not-dead misconceptions about this technology. It will also let you not step into the mud puddle of excessively optimistic promises by some vendors.

LEFT and RIGHT of SIEM

LEFT OF SIEM

What is to the left of SIEM? Mostly data collection. Data collection sounds conceptually simple, but operationally it is still very difficult for many organizations. It has been that for the entire quarter of a century of SIEM existence (the first SIM/SEM vendors launched in 1996, as a reminder).

To the left of SIEM lies a vast and unexplored — well, explored, but largely still unknown for many — land of data collection. Just as early SIM/SEM innovators struggled with collection [and then UEBAs did], innovators in 2022 struggle with it as well. When some vendors made some collection easy in a naïve way, it dramatically decreased the quality of their data (GIGO FTW!), reduced security insight and increased the amount of work needed to be done by their clients…

Today, we also collect a lot more context. Now, some prefer to have just logs collected and then left context collection to the clients, which is, in my opinion, a big mistake. To me, context is critical for detection and the more work we do to add context to logs early and automatically, the quicker path will be to a security outcome for a client. Better, richer context — better detection!

Practically, how do you get the left of the SIEM right? Frankly, for a broad enough set of log sources, you never get it completely right. What makes it easier is throwing away your SIEM that charges per gigabyte, and focusing on what you need to collect and retain, not what you can afford to (but, yes, there is a lot more to it than this). Still, focusing on collection (sources, messages, volumes, architectures, use cases, etc) is still be required to succeed.

And of course you will still have filtering — there’s always stuff that you need to not collect. Many years ago, I worked for a company that had a slogan of “collecting all logs” and while I can see the appeal of that simple message, it has become impractical in 2020s. There is just too much.

RIGHT OF SIEM

How do you get what’s right to the SIEM’s right? Well, let’s start with the elephant in the room: SIEM creates alerts and other signals and handling them (manually or automatically such as via SOAR [ha, tricked you :-)]) is the central part of what is to the right of SIEM.

Just as with collections, there’s no magic here, but response playbooks play a huge role. There is a lot already written on this (example), but seeing the SIEM as a “member” of a pipeline from sources to security outcomes is what allows one to get the right …well…right.

BTW, there may be more enrichment of alerts to the right of SIEM (such as via enrichment SOAR playbooks), not just logs to the left of it. Context for detection means additional context for alerts and even incidents (and hunting clues), not just logs.

Many things to the right of SIEM are about process. From alert triage to escalation and even tuning (that sort of connects right to left), the things to the right of SIEM that needs to be done right are practices and processes. This reminds us why success with SIEM means getting the processes right — and there is no other way.

Now a trick question that this discussion also touches: if what is on the left is truly broken, can you fix it from the right? This is debated, and I prefer to spend my time fixing both sides, if both are broken, rather than trying to fix the left from the right, but YMMV.

Practically, how do you get the right of the SIEM right? Know what you want your SIEM to produce, know your outcomes, and then work on the processes that alerts and signals flow into. BTW, this came as fuzzy advice, but hey … I am not paid to give advice anymore :-)

SO WHAT?

Some people are asking whether it’s more important to fix what to the left or what to the right. Recall again this passionate social media discussion about whether it’s better to tune the inputs or to use another tool (SOAR again) to clean the outputs from your SIEM?

To summarize, what has to be done right for success with SIEM needs to be done to the left and to the right of it.

Key tasks that you must succeed that sit on the left from SIEM: telemetry collection, context collection, scalable collection architecture, etc.

Key tasks that you must succeed with to the right of SIEM: alert triage, automation of some responses, collection of additional context needed for alert, etc

BTW, this may mean that whoever owns more of that chain gets better outcomes? First party data collection to quality native response action may win the battle in the long run … is this what XDR is about?

SHIFT LEFT?

As the discussion indicated, there are few chances to do that for real. To me, better log data would be a true “shift left” for SIEM. Better, more structured and more meaningful, and more security relevant log data will fix many problems, but I am not entirely hopeful here, especially after this experience back in the day….

Related posts:


Originally published at Medium.

Monday, January 10, 2022

New Paper: “Future Of The SOC: Process Consistency and Creativity: a Delicate Balance” (Paper 3 of…

Sorry, it took us a year (long story), but paper #3 in Deloitte/Google collaboration on SOC is finally out. Enjoy “Future Of The SOC…


Sorry, it took us a year (long story), but paper #3 in Deloitte/Google collaboration on SOC is finally out. Enjoy “Future Of The SOC: Process Consistency and Creativity: a Delicate Balance” [PDF].

If you missed them, the previous papers are:

My favorite quotes are below:

  • “This paper highlights ways to create a consistent set of core processes, yet still allow room for creativity within the process set for your SOC.” [A.C. — capturing this magic was our focus for the paper!]
  • “Strong thought out processes are sometimes (unfairly) seen as the most boring of the people, processes, and technology triad. But they are what differentiate mature organizations with capabilities from those with a collection of the latest shiny toys. ”
  • “The security community has become increasingly vocal in its belief that out-of-the-box use cases from nearly any vendor do not cut it anymore. This is not to disparage vendor’s products; it is simply an admission that anything intended to be globally applicable to thousands of customers in a constantly evolving threat landscape is bound to come up short for the most pressing threats. A marriage of out of the box use cases with robust and continuous engineering of detections and process automations is needed for the modern SOC to keep pace.
  • “True innovation can be scary for most security organizations because it is, to a certain extent, a commitment to organic and messy growth rather than measurable procedures. The challenge for a modern SOC leader is thus balancing the desire for consistency — backed by repeatable, predictable, and effective processes on one side — and the desire to harness human creativity, initiative, and perhaps even irrationality on the other side. ”
  • A highly functioning modern SOC, one that is able to anticipate and detect threats on their way in rather than on their way out, has likely attained that balance between consistency and creativity. But how is this balance achieved? The trick is to create an unconventional, but somehow harmonious mixture of consistent, repeatable processes and human, anarchic, and spontaneous creativity.“
  • “In some maturity models, the highest maturity level is called “Optimizing.” Perhaps this best captures the vision of how a good SOC should look. Ultimately, security professionals are not striving for consistent operations alone; they are aiming for this elusive maturity tier whereby the previous foundational levels are so well entrenched that the SOC can spend its time truly optimizing, in a living, ever-adapting model.
  • “The better road is to build consistency and grow through lower maturity levels, and then let creativity loose within the processes that are already built. ” [A.C. — for more details on how, see the paper!]
SOC Consistence vs Creativity visual from the paper

Enjoy! We are writing the final paper 4 as we speak :-)

Related blog posts:


Originally published at Medium.

Tuesday, December 21, 2021

Stealing More SRE Ideas for Your SOC

As we discussed in “Achieving Autonomic Security Operations: Reducing toil” (or it’s early version “Kill SOC Toil, Do SOC Eng”), your…


As we discussed in “Achieving Autonomic Security Operations: Reducing toil” (or it’s early version “Kill SOC Toil, Do SOC Eng”), your Security Operations Center (SOC) can learn a lot from what IT operations learned during the SRE revolution. In this post of the series, we plan to extract the lessons for your SOC centered on another SRE principle — evolving automation.

First, for many security operations teams automation in a SOC is about saving time by automating routine tasks. To me, this constitutes current conventional wisdom about automation in your SOC. However, in this post, we want to reveal a broader truth about automation in your security operations activities, drawing the lessons from the field of site reliability engineering (SRE). Naturally, I am not an SRE, but I feel that my analyst experience makes me qualified to translate or “port” the findings from their domain (SRE) to ours (SOC) — and to make new discoveries in this process too.

Security operations materials often point out that automation is a force multiplier, not magic. The SRE book says the same: “For SRE, automation is a force multiplier, not a panacea.”

However, the book also adds that “multiplying force does not naturally change the accuracy of where that force is applied.” This reminds us that automating a broken process often makes it more broken, but also that automating something that isn’t game-changing or systemic for a SOC would make you slightly better, if that.

By the way, this is why the most common starter SOAR playbook is about phishing, a major time-suck of many aspiring SOCs (I’ve heard one spent 40% of analyst time on phishing response and that was after the email security gateway did its work).

So people often point out that the value of automation is about saving time. However, both security operations center practitioners and SREs agree — consistency is also a big part of such value (“What exactly is the value of automation? Consistency!”). In fact, “automation provides more than just time saving, so it’s worth implementing in more cases than a simple time-expended versus time-saved calculation might suggest.” Think about it — it’s not only about saving time, scaling (“scale is an obvious motivation for automation”), but also consistency of what gets done whenever it needs to be done. By the way, how to make the security processes consistent yet allow for creativity, such as threat hunting? We will explore this in the next SOC paper in January.

Speed does come up a lot in SRE discussions of automation, after all “humans don’t usually react as fast as machines.” In the past, I largely implied (even in 2009) that sub-second speed matters little in security, especially in the day and age of 200+ day response timelines. Guess what? With ransomware, speed does matter. If you detect it via a ransom note, it won’t matter how good your SOC was …

To summarize, the main lesson from SRE is that “the factors of consistency, quickness, and reliability dominate most conversations about the trade-offs of performing automation.” These lessons work well when starting to make your SOC scale faster than the threats.

Further, I picked up a particular new insight from the SRE book, namely that automation separates the operation from an operator (“Decoupling operator from operation is very powerful.”). Why is it good? Glad you asked: “once you have encapsulated some task in automation, anyone can execute the task.” What does this solve? Some of the talent shortage problems in your SOC! This again gives us a chance to scale faster than the growth of threats and assets.

Here is another very useful reminder for your SOC from the world of SRE: “automatic systems also provide a platform.” What does it mean? That script you wrote is not a platform, even if it automates something. The way I think about it, the platform is a programmable entity, a base to develop other cool things. This means you have a chance to go for a more systematic automation of your current and future SOC activities.

Also, the SRE world delivers a very fun, slightly paradoxical, consequence: “A platform also centralizes mistakes. In other words, a bug fixed in the code will be fixed there once and forever” Think about it for a second! This is not about SOC being a great place to come and make mistakes … this is about the fact that you go to ONE place to look for mistakes, rather than chase them over 50 tools and 200 regional offices. Centralizing mistakes is awesome — and a new thought for me (and, I am assuming for many SOC practitioners as well).

Finally, “automation as a platform” leads us to metrics: “a platform can export metrics about its performance, or otherwise allow you to discover details about your process you didn’t know previously.” As you can guess, this delivers sizable — and positive! — implications for your SOC, given how hard security is to measure in general.

To my surprise, our SRE colleagues also pointed out a few negatives of automations. Now everybody likes to point out that problems with automation stem from automatic systems causing damage. This can happen both in the operations realm and of course in our beloved domain of cyber. Google SRE book describes beautifully horrible examples where many production systems at an, ahem, major tech company were deleted by automation, reimaged straight to demagnetized dust with enviable scale and effectiveness …

Now, what are the lessons? Here is a new idea as well: “Automation needs to be careful about relying on implicit “safety” signals.” What does that mean in a SOC? Well, a classic example would be blocking access based on badness, without checking for business criticality. We imply that it is safe to block access, but do we have an explicit “this machine is OK to auto-block” list? This is safe to shut down? This is safe to block access to? Using explicit safety signals for automation is a useful insight for me.

I have learned other challenges that are relevant to the world of security operations, which frankly I haven’t thought about. Many SOAR users complain that when the security tools change, EDR vendors change, APIs, logs change and other technologies evolve, their SOAR systems don’t always follow quickly enough. This is a well-known problem in the world of SRE: automation “being maintained separately from the core system therefore suffers from “bit rot,” i.e., not changing when the underlying systems change.”

Another lesson that we are starting to see in many security operations centers is that those automations that are infrequent, such as playbooks run upon seeing rare attack indicators are difficult to test. “Automation that is crucial but only executed at infrequent intervals and therefore difficult to test is often particularly fragile because of the extended feedback cycle.” It is easy to refine an efficient playbook that runs 10 times a day, but it’s much harder to run and refine a playbook that is supposed to help for a particular type of an advanced attack and may run twice a year, if that. How do we fix that? With more automation — test automation and simulations in this case.

Another great idea for your SOC is hiding deep inside the book. This has been characteristic of many leading security operations centers, and it has been discussed in many detection engineering articles, but it is definitely NOT common at many mainstream SOCs: “The most functional tools are usually written by those who use them.” This is why in our ASO workshops we explain that “SOC analysts” and “detection engineers” must go … and become one, or at least work together closely. “DevOps” your SOC!

We promised to discuss not just how to automate, but the evolution of automation. Here the news from the world of SRE is the most exciting. They chart a path of arriving to the autonomic system (that does not need extraneous automation) by starting from a manual approach and then evolving to automation. Here is what the book says:

  1. “Operator-triggered manual action (no automation)
  2. Operator-written, system-specific automation
  3. Externally maintained generic automation
  4. Internally maintained, system-specific automation
  5. Autonomous systems that need no human intervention”

While some of the above are not obviously related to SOC, let me try this:

SRE to SOC translations

That last step is interesting for sure. There is a lot of fun, thought-provoking stuff in SRE thinking related to “autonomous systems.” For example, they say that “software-based automation is superior to manual operation in most circumstances, better than either option is a higher-level system design requiring neither of them — an autonomous system.“ They further explain that non-autonomous “where automation replaces manual actions, and the manual actions are presumed to be always performable and available just as they were before.”

Finally, this is a security blog and so I was meaning to end it on a depressing note (compliance!). But then it turned out that SREs do “black humor” almost as well: “If we are engineering processes and solutions that are not automatable, we continue having to staff humans to maintain the system. If we have to staff humans to do the work, we are feeding the machines with the blood, sweat, and tears of human beings.” How is that for SRE noir?

Next we plan to dive into the SLOs of SREs and see what we can learn to make the SOC better!

A fully cooked version of this will be published on a Google Cloud blog, but I’d love some feedback and comments from either security or SRE side…

Thanks to Iman Ghanizada for ideas and brainstorming.

Related posts:


Originally published at Medium.

Monday, December 20, 2021

Anton’s Security Blog Quarterly Q4 2021

Sometimes great old blog posts are hard to find (especially on Medium) , so I decided to do a periodic list blog with my favorite posts of…


Sometimes great old blog posts are hard to find (especially on Medium) , so I decided to do a periodic list blog with my favorite posts of the past quarter or so.

Here is the next one. The posts below are ranked by lifetime views. This covers both Anton on Security and my posts from Google Cloud blog, and our Cloud Security Podcast too (subscribe).

Top 5 most popular posts of all times:

  1. “Security Correlation Then and Now: A Sad Truth About SIEM”
  2. “Can We Have “Detection as Code”?”
  3. “New Paper: “Future of the SOC: SOC People — Skills, Not Tiers”
  4. “Beware: Clown-grade SOCs Still Abound””
  5. “Revisiting the Visibility Triad for 2020”

Top 5 posts with the most Medium fans:

  1. “Security Correlation Then and Now: A Sad Truth About SIEM”
  2. “Beware: Clown-grade SOCs Still Abound”
  3. “Can We Have “Detection as Code”?”
  4. “Why Is Threat Detection Hard?”
  5. “A SOC Tried To Detect Threats in the Cloud … You Won’t Believe What Happened Next”

Top 5 Cloud Security Podcast by Google episodes:

  1. Episode 1“Confidentially Speaking”
  2. Episode 2 “Data Security in the Cloud”
  3. Episode 17 “Modern Threat Detection at Google”
  4. Episode 8 “Zero Trust: Fast Forward from 2010 to 2021”
  5. Episode 27 “The Mysteries of Detection Engineering: Revealed!”

Random fun new posts:

  1. “SOC Technology Failures — Do They Matter?”
  2. “Kill SOC Toil, Do SOC Eng”
  3. “Anton and The Great XDR Debate, Part 1”

Fun posts by topic.

Security operations / detection & response:

Data security:

Cloud security:

Enjoy!

Previous posts in this series:


Originally published at Medium.

Friday, December 10, 2021

SOC Technology Failures — Do They Matter?

Most failed Security Operations Centers (SOCs) that I’ve seen have not failed due to a technology failure. Lack of executive commitment…


img src: https://flic.kr/p/dwWHw5

Most failed Security Operations Centers (SOCs) that I’ve seen have not failed due to a technology failure. Lack of executive commitment, process breakdowns, ineffective workforces (often a result from poor management and lack of commitment … again) and talent shortages have killed more SOCs than any and all technology failures.

Example SOC Troubles from some presentation :-)

As we are working on the next SOC paper jointly with Deloitte (paper 1, paper 2, paper 3 coming out really soon), we came across the need to review some of the current technology challenges in the SOC. Hence this blog was born.

BTW, if somebody wakes me up at 3:00 a.m. and says “Anton, what is the top reason why a security operation center may fail?” I would name the loss of executive commitment. I have seen too many SOCs that decayed over time as management lost interest in their excellence, then in their performance and finally in their existence… Some of the noted breaches in the last decade can be traced to a SOC that was developed, refined and improved and then left to deteriorate (naïve outsourcing often played a role of a final nail in the coffin). On the flipside, there is nothing better to revive the presence of the SOC other than a major security incident — the trick is to keep the momentum going for years afterwards…

But I digress. Let’s stick to mostly technology focused failures. An astute reader will notice that in the list below, some of the purported technology failures are really process failures in disguise. I apologize for this in advance :-)

Scaling failures — Often when you get to POC an impressive new tool, the ability to scale is not truly understood and battle-tested until you onboard the tool into your workflow. There can be several dimensions to scaling from the ability to consume data, process data, store data, and make sense of all your unique data types at the scale you need it to be. Even the way you use the tool may cause unseen bottlenecks and affect your ability to scale. It’s far too often that vendors showcase a product’s abilities in it’s best and perfect use case, without little regard to scale. In other instances, the tool may have strong engineering behind it and truly have an ability to scale to many use cases, but a terrible user experience that makes everyone dread using it with the data volumes at hand (think drop-down menus with 2300 entries…) Or, the tool works, but only as long as you down-scope your collection efforts to below your actual security needs. Finally, the tool may “scale physically, but not economically” i.e. it will run at scale you need, but nobody can realistically afford it …

Tool deployed and then not operationalized sounds like a process failure, or a people failure. I lamented on this back in 2012, and this affliction has not truly subsided. However, why are some tools sitting unused in those boxes while others develop an active and passionate user bases? You don’t think it can be about the tool at all? Perhaps the tool vendor made some incorrect assumptions about how their technology is really used in the real world? Or, overestimated it’s value a bit for a particular type of client? Or a client may have thought that deploying the tool and turning it on was “self-service” — only to realize that they should have paid for the consulting partner to make it work.

In other cases, the tool “does runs, but does not work” (my name for failure to design the tool to be usable in real-world environments). There are also piles of tools that are deployed and used in production — but only 10% of the capabilities are utilized. Some technologies seem valuable, but are such a burden to maintain and use in real life that it is practically impossible. SOC should not spend time / resources managing such technologies. If your key SOC technology (whether you call it a SIEM or not) is that hard to manage, toss it, buy SaaS-based technology.

All in all, SOCs that have to manage too many tools suffer. A pile of boxes (yes, and cloud services) that need care and feeding, tuning, refining can overwhelm even a large team. Buy what you would use, and use what brings value!

Shiny new tool syndrome is still rampant in some SOCs. A new CISO comes in, tries to champion the implementation of a new tool, the CISO is gone after a short amount of time — like most CISOs, and then a new CISO comes in and tries it all over again. And tools multiply, getting less shiny day by day. Or, the SOC team champions a build-first strategy in areas where it is better to buy — as they figure out that building scalable solutions with open source tooling becomes way more challenging than initially thought. Yes, DIY SOC tools fail as well.

Data collection failures still plague many SOCs. Now, again, one can also blame this on people and processes (especially, those people in IT who just didn’t give us the data). However, in many cases it is in fact the tools (such as when a pre-cloud security monitoring tool is aimed at the cloud). Data collection does not stop at getting the data, because things like the lack of uniform data model, challenges with getting value out of raw data that is not enriched belong in the same category.

One sided visibility stack is definitely a tool challenge as well. A SOC that only uses a SIEM, only uses an EDR, or (if you are crazy) only uses an NDR is missing out. As I noted in the SOC visibility triad discussion (2020 refresh), there is a decent chance that in the near a future a SOC that uses a 2015-style triad of SIEM+NDR+EDR is also missing out, such as on the application security telemetry, as organizations develop more security use cases for observability data.

Somewhat related, old tools that don’t cover new environments sounds like a tool challenge, and the above cloud example works here as well. Frankly, many traditional SOCs suffer with cloud, with containers and with other modern IT technologies and environments (old example). You don’t need to buy the whole lot (CWPP, CSPM, CASB, SSPM, CNAPP, etc), but you do need to be mindful of public cloud visibility gaps in your SOC.

Along the similar line, tools that promise “a single pane of glass” usually don’t deliver that. Don’t wait for a vendor to invent a “single pane of glass”, it either won’t come, or it will come out of the box “pre-broken.” I am not sure if the mesh is the answer either. Note that integrated tools are hard to adopt if you already have siloed tools that are at least partially successful (will you ). Would you buy an XDR that includes an EDR if you are happy with your different EDR?

Automation is sometimes more work than value — even with SOAR. Again, this may be seen as a people challenge (“hey, you just don’t have enough security developers in your SOC ….oh wait … you have none”). In real life, automation is of course a benefit, but it does require work to deliver. Indeed, this reads like another people/process challenge, but perhaps tools can eventually deliver real — not excessively optimistic — low code/no code security automation?

Thanks to Iman Ghanizada for the ideas and contributions to this post!

Related posts:


Originally published at Medium.

Dr Anton Chuvakin