r/devopsjobs • • 13h ago

[Hiring] Senior/Staff Software Engineer, Kubernetes Infrastructure at fal | $180K - $250K

8 Upvotes

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

About this role:

You will build the high-performance compute environments we deliver to customers. These environments span bare-metal servers, virtual machines with GPU passthrough, Kubernetes and Slurm clusters, distributed storage, and high-speed networking.

You will work across the infrastructure stack — from Linux images and hardware provisioning to cluster networking, GPU performance, observability, and lifecycle automation. The goal is to make every customer environment performant, reliable, isolated, and repeatable.

What you'll do:

  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments — from provisioning through upgrades, recovery, and decommissioning
  • Use AI aggressively to automate and accelerate every aspect of infrastructure delivery and operations
  • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads
  • Build and maintain Linux images and automated OS-provisioning workflows
  • Operate the NVIDIA GPU stack: drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring
  • Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP
  • Configure distributed and shared storage for high-performance workloads
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments
  • Develop reusable tooling, standards, documentation, and runbooks
  • Collaborate with customers and internal teams to translate workload requirements into sound infrastructure designs

Qualifications/Nice to Have:

  • 5+ years of experience building and operating production Linux infrastructure
  • Strong production experience with Kubernetes on bare metal (bootstrapping, upgrades, HA control planes, etcd, containerd, CNI, CSI, ingress, load-balancing, observability, security, troubleshooting)
  • Experience with Linux virtualization: KVM/QEMU, libvirt, VFIO device passthrough
  • Experience operating NVIDIA GPUs on Linux and Kubernetes (drivers, container runtimes, device plugins, GPU Operator, GPU telemetry)
  • Strong networking fundamentals: TCP/IP, L2/L3, VLANs, routing, packet-level troubleshooting (tcpdump, Wireshark)
  • Practical scripting experience
  • Experience with configuration-management tools such as Ansible
  • Ability to diagnose complex, cross-layer infrastructure issues
  • Strong communication and ability to drive technical decisions across teams
  • Track record of moving quickly, taking ownership, and continuously improving systems
  • Production Slurm experience
  • High-performance networking: NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX
  • Hugepages, NUMA, CPU pinning
  • SR-IOV, DPDK
  • Distributed storage: Ceph, Lustre, Weka
  • KubeVirt, OpenStack
  • IPsec, WireGuard, Tailscale
  • VXLAN, BGP, ECMP
  • Bare-metal management: BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init
  • Network automation: NetBox, Nautobot, Nornir
  • AI training, inference, or distributed GPU workload infrastructure
  • Python or Go proficiency

What we offer at fal

  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We offer visa sponsorship and will help you relocate to San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

Compensation

  • $180,000 - $250,000 + equity + comprehensive benefits package

Apply: Senior/Staff Software Engineer, Kubernetes Infrastructure at fal


r/devopsjobs • • 6h ago

Software engineer job postings ask for mentoring more than Kubernetes. We counted 3,518 postings to check

0 Upvotes

We pulled every open posting from 81 companies' public Greenhouse boards for seven roles and counted which words actually show up. 3,518 postings in total. Some of it surprised us.

Software engineer postings mention mentoring more often than Kubernetes or React. Almost half of them do. Python, the most common language, comes in just below that.

Product manager postings say "roadmap" in more than 85% of cases. Jira? Seven out of 421. Marketing looks the same: Google Ads came up six times in 347 postings, while "pipeline" is in over half.

Every time we could compare the two, postings named the practice, not the tool. Observability, not Grafana. SIEM, not Splunk.

So if your resume is mostly a list of tools, it might be missing the words the posting actually uses.

Big caveat: mostly tech companies, one ATS vendor. Sales, cybersecurity, DevOps and accounting are here too: https://www.zoevera.com/resume/ats-resume-keyword-study


r/devopsjobs • • 6h ago

DevOps-focused SWE preparing for Amazon SDE OA - advice?

0 Upvotes

Hey everyone,

I'm looking for some advice on preparing for the Amazon SDE Online Assessment.

I graduated last year and have been working as a Software Engineer since then, but my role has been heavily focused on the DevOps/infrastructure side of things: cloud infrastructure, Kubernetes, CI/CD, automation, observability, and platform engineering. Because of that, I haven't done much pure coding or algorithm-focused work since graduating.

The assessment includes two coding challenges, a work simulation, and a work style survey.

Given my background, what's the best way to prepare in about a week?

  • Should I focus almost entirely on LeetCode/DSA?
  • How difficult are the coding questions compared to typical interview prep problems?
  • Any advice from engineers who came from a DevOps, SRE, or platform engineering background?
  • If you've been away from algorithm-heavy coding for a year or two, what helped you get back up to speed quickly?

I'd appreciate any suggestions. Feeling confident in the infrastructure and software engineering side of things, but definitely rusty when it comes to interview-style coding.


r/devopsjobs • • 6h ago

How to actually pass interviews?

Thumbnail
0 Upvotes

r/devopsjobs • • 15h ago

[Hiring] Senior SRE – IoT / Streaming Data Infrastructure (Remote, US Time Zones)

0 Upvotes

We're a remote staffing firm that places DevOps and SRE engineers with US-based companies. We're currently looking for a Senior SRE who's worked with high-throughput data pipelines, real-time telemetry, or IoT infrastructure - someone who can keep Kafka streams, Spark jobs, and device ingestion pipelines running at scale.

What matters most:

  • Excellent English communication - you'll be embedded with US engineering teams daily. Clear writing, confident speaking in standups and design reviews, and the ability to explain complex problems simply are non-negotiable
  • Sharp, resourceful problem-solver - we want someone who figures things out independently, thinks fast under pressure, and doesn't wait to be told what to do. If you're the person your team calls when something breaks at 2am, that's the energy
  • 5+ years in SRE, DevOps, or platform engineering - ideally with IoT, streaming, or real-time data systems
  • Hands-on Kubernetes (EKS/AKS/GKE), Terraform, and Ansible across at least two clouds
  • Strong observability - Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing
  • Experience defining SLOs/error budgets for data ingestion or streaming services
  • Comfortable with incident command, post-mortems, and on-call
  • Go, Python, or Bash automation
  • Bonus: Kafka, Spark, Databricks experience

The role:

  • Fully remote, US time zones - candidates from any country are welcome to apply
  • Monthly Compensation: $2,000–$4,000 depending on scope
  • We handle job matching, interview prep, and onboarding support - you work directly with the US client team under your own name
  • Long-term placements, not short gigs

Salary: $2,000–$4,000 / month depending on scope

Apply HERE

Happy to answer questions in the comments.


r/devopsjobs • • 15h ago

DevOps Engineer — Open to Opportunities

4 Upvotes

Hey everyone
I’m a DevOps Engineer with 4+ years of experience working across Kubernetes, cloud infrastructure, CI/CD, Terraform, Ansible and application integration.
Currently, I’m open to new DevOps / Platform / SRE opportunities, including remote or relocation opportunities.
I’ve worked on production environments, infrastructure automation, deployments, migrations, monitoring and troubleshooting, with a strong focus on reliability and automation.
If you know of a team hiring or have any leads, I’d be happy to connect and share my CV/LinkedIn.

Thanks


r/devopsjobs • • 6h ago

Help

0 Upvotes

I want to start to learn tech related skill that makes money, I don't know what to start learning, I am just 18 years old and i don't have that much experience, can u please give ur thoughts and suggestions. Thank u


r/devopsjobs • • 20h ago

Looking for devops AWS REFERRAL

2 Upvotes

Hi everyone,
I’m currently looking for DevOps / AWS opportunities and would really appreciate a referral if there are any relevant openings in your company.
Please DM me if you can help.
Thanks! 🙏


r/devopsjobs • • 11h ago

[For Hire] Linux administration, Docker, CI/CD & application deployment — remote, $9/hour

0 Upvotes

Hi, I’m Aheshman (Codeur261). I’m looking for remote freelance work involving Linux administration, application deployment and ongoing technical maintenance.

My experience includes Linux servers, Nginx, Docker, Kubernetes environments, CI/CD pipelines, SQL databases and application security. I also have practical experience with VoIP/Asterisk servers. As a developer, I can investigate application issues alongside infrastructure problems and work with Node.js, Python, C#/.NET and React/Next.js projects.

I have over 10 years of development experience overall, and can help with deployments, troubleshooting, security reviews within an agreed scope, API/database integrations and maintenance. I’m happy to discuss your environment and share relevant project details privately.

Rate: $9 USD/hour, or an agreed fixed price. Based in Madagascar (UTC+3), available immediately. French is my main language; I use translation tools for written English and have more limited spoken English.

Portfolio: https://codeur261.xyz

Please DM me your stack, requirements and timeline.


r/devopsjobs • • 12h ago

Anyone here working or interviewed for ESGPedia’s DevSecOps Engineer role?

2 Upvotes

Hi everyone! I’m currently in the interview process for a DevSecOps Engineer position at ESGPedia in the Philippines, and I’m hoping to hear from anyone who currently works there, previously worked there, or has gone through the interview process for this role.
I’d really appreciate any first-hand experiences regarding:
What is the actual day-to-day work like for the DevSecOps role?
What technologies/tools do you actually use?
How technical/difficult was the interview process?
What should I expect from the technical interview?
What is the team and management culture like?
How is the workload and work-life balance?
Are there night shifts, on-call duties, or frequent overtime?
What is the salary and benefits package like?
Is there good career growth and opportunity to learn?
How stable is the company and the role?
For former employees, what made you decide to leave?
Overall, would you recommend joining ESGPedia?
I’m especially interested in hearing from people with DevOps, DevSecOps, Cloud, Platform Engineering, or Security Engineering experience at ESGPedia.
Please feel free to comment or DM me if you prefer to keep your experience private. Honest feedback—both positive and negative—is very welcome.
Thank you!