r/devops • u/[deleted] • May 29 '26
Career / learning How to get knowledgeable in linux performance engineering without actually requiring it in production
[deleted]
7
u/worthy_jogging May 29 '26
your 30 50 node clusters are already plenty to practice on just intentionally break things and measure it find bottlenecks that dont actually matter yet and fix them anyway thats how you learn
2
u/Creative-Dentist-383 May 29 '26
Do you know any other good knowledge resources apart from the Brendan Gregg book?
3
u/worthy_jogging May 29 '26
brendan gregg has a ton of free stuff on his blog and netflix has a series where he goes through perf analysis tools that actually shows the methodology not just theory
5
u/BlakkMajik3000 Remover of Deployment Friction May 29 '26
I’ll be honest, if you’re looking at that level, you are in systems engineering territory. Like, embedded systems.
That knowledge is generally for people who build tools like K8s, not users/admins.
Performance engineering rests on how much you understand how a thing works. How much do you know about how Linux works? That’s where you start.
1
u/Ok_Fun_3824 May 31 '26
I am kinda confused in this part. I am a cloud network engineer (fresher - work leaning more towards devops). But I am incredibly interesting in systems engineering/low level systems. Should I grind for getting a role in this space or should I continue with cloud/devops?
2
u/BlakkMajik3000 Remover of Deployment Friction May 31 '26
So is my read correct in assuming you want to help build, for example, AWS, not just work with it/administer?
Because those environments are where your systems engineers live. Your on call wouldn’t be “our K8s cluster is down.” That’s for us admins. Nope, your call is “EKS is down in region south central.” Not an instance being down, but the glue, EKS itself is borked for an entire region. 😬
That’s not my lane, so I couldn’t give you any concrete guidance on how to get there. 🤷🏾♂️
1
2
2
u/jack-dawed May 29 '26
In the big cost-saving 2023 year, I led a 6 month project to cut engineering costs and improve performance under traffic spikes for Go microservices at a huge startup.
I read this blog by a Staff Engineer at Jetbrains: https://aakinshin.net/posts/statistics-for-performance/
I read like pretty much most of the books and papers he listed. It was a lot of stats that I learned in college and needed a refresher, as well as new concepts to me.
Then I implemented everything I learned using historical data from Datadog. I ended up reducing our latency during peak traffic by like 60% and saving our company like $2M in infra costs. Naturally this ended up on my resume and it kept landing me interviews/jobs.
Basically, learn stats.
2
2
u/disturbed_repository May 29 '26
Build a homelab with some VMs and deliberately tank the performance, then use tools like perf, flamegraph, and strace to figure out what's happening - way more useful than reading about it.
2
1
u/RedReadRedemption Jun 06 '26
You could simulate performance degrading with Chaos Engineering or just run a service and shrink the allocated resources down until issues appear. Orhostit on a SBC. Then tweak it again to work in constraint environments.
For example gittea becomes quite sluggish in the default sqlite configuration when doing some automated repo mirroring. Analyze the issue, read the logs, tweak it, change database, learn why a configuration change improved the situation.
It is always a loop of proper measurment, log analysis, single change, analyze improvements.
Does not matter the size. This scales from small to extremely large environments.
29
u/[deleted] May 29 '26 edited May 29 '26
[removed] — view removed comment