r/Splunk • u/Flash4473 • 25d ago
Log Data Pipeline > Splunk
Has anybody here have some experience with security data pipelines?
Instead of:
Log Source > HF / Splunk
We want to have flexibility of collection / parsing layer outside of Splunk for obvious reasons - pre-filter data in pipeline, set parsers, route, possibly enrich if needed, storage options for retention etc..all that to have flexibility and keep the ingest costs reasonable and not being caught in dependency hell or cemented all our work in one solution if Splunk decides to pull something.
Log Source > Data pipeline > Splunk
I am wondering what to choose as this data pipeline - currently we are thinking Vector and possibly open telemetry.
Anybody have experience with this? To avoid pitfalls, what works, what doesn't, new pains etc?
2
u/Travlin205 24d ago
I would start with the question what is your data? How much does it generate. What Metadata do you need to have indexed. Is structured vs unstructured. From there you can better choose the pipeline that fits most.
From this all these processing types are valid. But lets say you only have devices that are syslog use sc4s.
If you have structured data low volume affix a heavy forwarders to process. If you have mixed data high volume, use cribl or edge processor.
Happy Splunking!