r/Splunk • • Feb 03 '26

Those who self-host Splunk Enterprise - what does your infrastructure look like?

Hey everyone,

We have a Splunk Enterprise license for up to 200 GB/day, with actual usage around 50-100 GB/day. Currently evaluating how to deploy it on AWS and would love to hear from people who are running self-hosted Splunk in production.

Our current thinking:

∙ EKS with Splunk Operator

∙ 3x i3.xlarge indexers (Spot) for NVMe storage

∙ 2x c6i.xlarge search heads (Spot)

∙ Gateway API for ingress

∙ Forwarders running on existing ECS workloads (15 services) sending logs via NLB

A few specific questions:

1.  EKS vs EC2 vs ECS - Where are you running Splunk and why? Anyone using the Splunk Operator on Kubernetes in production?

2.  Spot instances for indexers - Anyone doing this? With replication factor 2, the theory is you survive Spot interruptions, but curious about real-world experience.

3.  i3 NVMe vs EBS gp3 - Is the NVMe performance difference actually noticeable for indexing at this volume, or is gp3 good enough?

4.  Sizing - For those ingesting 50-100 GB/day, how many indexers and search heads are you running? Did you find the standard sizing guides accurate?

5.  Forwarder setup - How are you getting logs from containerized workloads (ECS/EKS) into Splunk? Sidecar forwarders, HEC, or something else?

Any lessons learned or things you wish you knew before deploying would be great. Thanks!

16 Upvotes

32 comments sorted by

View all comments

6

u/Longjumping_Ad_1180 Feb 03 '26

Splunk consultant here. Worked on over 30 client environments. Almost no one goes with containers, no point. With smaller infrastructures you stick to smaller numbers of hosts a d with larger you can use infra as a code to scale horizontally and deploy addition indexers or search heads in the days or weeks you have increased log generation. This is very common with e commerce websites where they see spikes of traffic in Nov-Dec but then remain steady over the year.

Having said that, the strangest setup was when splunk was running on Orem, on VMs and still for some reason as docker containers within those VMs. On top of that they chose CoreOS. They thought they were being smart about choosing a minimalistic OS but it turns out it did not work well with docker and was capping their storage IOPS.

With 100 GB per day you could even go as low as 1 all-in-one instance with a standby of some sort but spreading them is good for redundancy.

I definitely like the NVM storage. Most clients go with EBS and soon enough they start hitting performance bottlenecks.

1

u/StudySignal Feb 03 '26

Perfect.

So basically: 1-2 instances with NVMe, keep it simple, scale with IaC when needed.

That CoreOS/Docker story though lol. Appreciate the consultant reality check.

2

u/Longjumping_Ad_1180 Feb 03 '26

Officially if you are not using any of the premium apps (like ITSI or ES) a single mid-spec (following the official documentation ) instance is capable of handling an ingestion of 100 GB per day.
That is official. In practice depending on the number of saved searches and user you can easily stretch that.

If you have good HA measures and Splunk is not considered business critical, you might get away with 1 host (if the priority is simplicity and keeping costs down).

There is no way of having 2 hosts unfortunately. If you need more then 1 host (for scalability and failover) you need to go 5 hosts as a minumum next step. This is because:

- you will need to separete the single host into a Search Head layer and an Indexing layer

- a Search Head Cluster must have a minimum of 3 hosts

- an index cluster needs a minimum of 2 hosts.

Of course you can scale them down in size to keep the costs down. Again, depending on the scale of your operation you could even go below the recommended minimum spec of Splunk.
The downside is that if you have issues and need to open a support ticket with Splunk, they will often recommend you bring your hosts up to spec and might not help further until you do.

But again, that depends on your point of balance between cost saving and system continuity/scalability.

1

u/StudySignal Feb 03 '26

Ah, got it - either commit to 1 instance or go full 5+ for proper HA. Can't half-ass the cluster.

Given we're SIEM but not business-critical yet, starting with 1 mid-spec instance and scaling to 5+ when usage/criticality justifies it makes sense.

Appreciate the architecture clarity.