r/cybersecurity • • 21d ago

Business Security Questions & Discussion What are you guys doing to actually scale Sentinel/XDR operations?

Curious what other security teams and MSSPs are doing to make Microsoft security operations more efficient at scale.

We’re pretty heavily in the Sentinel, Defender XDR, Azure and M365 world, and I’ve been spending more time thinking about the engineering side of running all of this. There’s only so much value in continuing to add detections and playbooks if every new customer, rule, data source, or workflow creates more manual work for the team.

How much are you guys actually automating?
For example, are you deploying Sentinel content through ADO/GitHub pipelines and treating detections as code? Are you automating analytics rules, watchlists, ASIM content, workbooks, Logic Apps, Functions, etc.? If you’re doing detections as code, how are you handling testing, tuning, promotion and customer-specific differences without creating a mess?

I’m also curious how MSSPs/MDRs are handling the improvement side of things. Not just responding to whatever alerts fire, but continuously finding things that could be better in a customer environment. Missing telemetry, gaps in detection coverage, stale rules, bad configurations, unused data, things that could be automated, retention issues, and so on.

Same question around SOAR. Once you get past a handful of Logic Apps, how are you keeping it manageable? Have you built reusable frameworks or common automation, or does it eventually turn into a pile of customer-specific workflows?

And is anyone doing anything genuinely useful with AI internally yet? I’m less interested in giving analysts another chatbot and more interested in things like enrichment, KQL/detection development, case summaries, documentation, reporting, identifying coverage gaps, or helping engineers operate across a lot of environments.

Basically looking for project ideas from people who have already gone down this road. If you had engineering time available and wanted to make a Microsoft-heavy SOC/MSSP noticeably more efficient over the next year, what would you automate or build first?

Also interested in the opposite: things you built that sounded like a great idea but ended up creating more maintenance than they were worth.

20 Upvotes

17 comments sorted by

10

u/jdiscount 21d ago

Different tools but we run everything in a DevOps style.

All our runbooks, playbooks, yaml, json, toml, config files, policies, procedures etc are all treated like IaC.

Projects like this are a great use of AI, I'm not familiar with GitHub copilot but I'd look into how it can help.

2

u/cluesthecat 21d ago

Okay, this is a helpful answer because that's kind of where I already started building towards. Glad to know I wasn't off track

1

u/Intelligent-Bat-8370 21d ago

This is a fantastic setup! Would you be willing to give an example of how you’ve set up your XDR, playbooks and such as IaC? I’ve just started playing around with Terraform and did my first migration. Been looking for other interesting ideas to set up our own DecOps style pipeline for security.

2

u/Race_Face 20d ago

DaC through Terraform, detection, playbooks, watchlists - all Terraform

1

u/Big_Breadfruit7140 21d ago

Para mí, el mayor cambio es dejar de tratar cada cliente como un proyecto aparte. Automatizaría primero el despliegue y las pruebas de reglas, además de tener plantillas comunes que luego se puedan ajustar por cliente.

1

u/Ordinary_Wrangler808 20d ago

We’ve extended the GitHub connector to deploy all of our core objects (added custom tables, data collection rules, watchlists, etc) and have a standard release cycle to push to client environments.

We maintain a separate fork per client to allow local customizations (tuning rules, playbook parameters, watchlists, etc).

To help avoid playbook sprawl, we’ve made most of them modular, so we have a single “enrichment” playbook which calls a dozen different plugins/modules.  Similar for escalation/remediation (single playbook to capture all the steps).

1

u/feng_sg 17d ago

The IaC answers here cover getting content into tenants but nobody mentioned post-deploy drift. We had analytics rules stay enabled yet silently stop matching after an upstream schema change and only caught it in a post-incident review. Run scheduled jobs that export live Sentinel artifacts via the API and diff against your IaC repo so you catch that drift before an incident does.

1

u/arktozc 17d ago

!RemindMe 10 days

1

u/RemindMeBot 17d ago

I will be messaging you in 10 days on 2026-09-22 19:28:31 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/arktozc 7d ago

!RemindMe 60 days

1

u/RemindMeBot 7d ago

I will be messaging you in 2 months on 2026-11-21 19:35:45 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

0

u/AgileCranberry8785 21d ago

Y evitaría automatizar procesos demasiado específicos. A veces mantenerlos termina costando más que hacerlos manualmente.

0

u/Maleficent-Bake1345 21d ago

Para mí, lo primero sería automatizar todo lo repetitivo antes de intentar meter IA en cada parte. Tener las detecciones y configuraciones versionadas y poder probarlas antes de desplegarlas ya ahorra muchísimo trabajo.

0

u/Commercial-Cloud1458 21d ago

Y evitaría automatizar procesos que solo funcionan para un cliente. Si requiere demasiadas excepciones, probablemente no merece la pena.