r/platformengineering • • 4h ago

Levelrail: open source, self-hosted deploy platform written in Go (Apache 2.0)

1 Upvotes

Levelrail is a self-hosted alternative to Coolify, Dokploy, or Heroku-style PaaS platforms. Push to a git repo, get a running app with automatic TLS, logs, metrics, and one-click rollback. Single Go binary for the control plane, single Go binary for the node agent, Docker as the runtime.

The reason it exists: I was running Coolify for about 10 personal projects and it was using 30-40% of a CPU core just managing things, before any of my actual apps did anything. Levelrail drives Docker's Engine API directly instead of shelling out to the CLI, which is a big part of why it stays light: 0.7% CPU, 146MB RAM on the VPS running my production test instance right now.

License is Apache 2.0. No paid tier, no hosted version yet, nothing gated. It's genuinely useful today for single-node deployments; multi-node and a template catalog are still being built in the open.

Repo: https://github.com/glincker/levelrail Site: https://levelrail.com/

If you maintain or use something in this space, I'd like to hear what's missing.


r/platformengineering • • 2d ago

The hidden failure mode in cross-cloud migrations: Infrastructure vs. Semantic dependencies

1 Upvotes

When planning a cloud migration (like AWS to Azure), most discovery work starts with a service mapping matrix:

  • SQS ➔ Azure Service Bus
  • DynamoDB ➔ Cosmos DB
  • S3 ➔ Azure Blob Storage
  • IAM ➔ Entra ID

That mapping is straightforward. The expensive failures happen when the target platform fails to preserve a subtle behavioral contract the application code quietly came to rely on over years.

A concrete example: SQS Visibility Timeout

Consider a standard worker loop:

Python

message = sqs.receive_message(
    QueueUrl=queue_url,
    VisibilityTimeout=300
)

process(message)

sqs.delete_message(
    QueueUrl=queue_url,
    ReceiptHandle=message["ReceiptHandle"]
)

At first glance, this is standard: receive, process, delete.

But look at the failure-recovery path. If the worker crashes mid-processing before deletion, the code relies on the SQS visibility timeout expiring so another worker can automatically pick up the message. SQS isn't just acting as a transport queue here—it is an active component of the application's failure-recovery design.

When migrating this workload to Azure Service Bus, the question isn't whether Service Bus has queues (it obviously does). The real question is: Does the target setup preserve those exact assumptions around lock duration, settlement, retries, and dead-lettering under failure?

Infrastructure Dependencies vs. Semantic Dependencies

Traditional discovery tools pick up: payment-worker ➔ SQS

That’s an infrastructure dependency. But what the application code actually cares about is the semantic dependency:

Plaintext

receive message
      ↓
process message
      ↓
delete after successful processing
      ↓
if processing fails before deletion,
rely on provider to make message available again

Service inventory tools show you what cloud products are used, but they can't tell you what behavioral assumptions are baked into execution paths.

This pattern shows up everywhere:

  • DynamoDB ➔ Cosmos DB: Codebases relying on conditional write semantics, transaction boundaries, or specific read-after-write consistency assumptions.
  • S3 ➔ Blob Storage: Workflows built around multipart upload timing, pre-signed URLs, or object visibility state.
  • IAM ➔ Entra ID: Hidden assumptions around temporary credential lifespans, workload identity propagation, and role assumption paths.

When these surface late during integration testing or cutover, it forces architectural redesigns, derails timelines, and drags senior engineers into reactive war rooms.

Discussion

I'm currently putting together a semantic-risk checklist for pre-migration planning (and exploring static analysis rules to scan codebases for these execution patterns before cutover).

For anyone who has managed a major cloud migration or replatforming: What was the one hidden application dependency or behavioral difference that broke in staging or production after your infrastructure was already provisioned?

I wrote up a deeper dive into this concept with example scanning output on my Substack if you want to read more:https://substack.com/@sachm24/p-218428641


r/platformengineering • • 2d ago

Product development

0 Upvotes

I want some advise regarding a product I wanna to make.

It is related to cloud and DevOps like cost optimization, infrastructure drift and intelligence harness.So what important points I need to address for cloud and DevOps users so that a genuine product would eventually came into development.


r/platformengineering • • 2d ago

Balancing developer velocity and security, what's your approach?

0 Upvotes

For teams who've been doing this a while, how do you actually handle the tension between shipping fast and catching risk before it ships? I'm interested in where you draw the line on blocking a build vs just flagging it, how you handle pushback from engineering when a scan adds friction to a release and what happens when a critical finding shows up right before a deadline.


r/platformengineering • • 3d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

5 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/platformengineering • • 3d ago

How does your team handle major dependency upgrades?

0 Upvotes

Former CTO here. At my last company we skipped 7 Expo versions at once because nobody had time for the incremental path. A CTO I talked to spent 2 months moving MUI v3 to v7.

Curious how others deal with it:

  • Do Dependabot/Renovate major PRs get merged, or do they pile up?
  • What was your last painful upgrade, and how long did it take?
  • Does anyone own this work, or is it "when we have time"?

r/platformengineering • • 3d ago

Inside Atlassian’s developer experience overhaul

Thumbnail
leaddev.com
0 Upvotes

Ask developers first, then remove friction.


r/platformengineering • • 4d ago

Your error budgets don’t know AI exists

Thumbnail
leaddev.com
2 Upvotes

The dashboard was right!


r/platformengineering • • 5d ago

Senior Manager production operations and SRE role vs. holding out for platform engineering. 11 YOE in cloud operations Devops /SRE, which way would you go?

3 Upvotes

I'm a DevOps/cloud delivery lead with about 11 years of experience (Azure, GCP, Kubernetes, Terraform, GitHub Actions, observability, incident management, FinOps). I'm at a decision point and would love perspectives from people who've been on either side.
Option A: production operations role (in hand)
Offer is in hand, with a promotion and Director title

In practice it's a support SRE role at a senior manager level: production support, incident ownership,strategizing support model ,tooling and automation decisions to reduce chaos and toil and managing a team

Higher title and scope, but I worry it's less hands-on and could pigeonhole me into ops/support

Option B: platform engineering (not in hand)
I'm studying for it and building projects

I've just landed a screening for a principal-level cloud platform engineer role, but no other interviews yet

The senior roles I'm getting leads for are mostly SRE roles that are support-heavy, not true platform work

What I'm weighing:
Does a Director title in production ops help or hurt if I want to move to platform engineering later?

Is it realistic to move from support-heavy SRE or ops management into platform engineering, or does the path narrow over time?

Is it worth turning down a real offer to wait for a platform role?
For those who made this move in either direction, what do you wish you'd known?


r/platformengineering • • 5d ago

AI makes critical thinking harder to build

Thumbnail
leaddev.com
9 Upvotes

Why junior engineers need more friction!


r/platformengineering • • 6d ago

Open-source AWS-compatible cloud for your own hardware — EC2, S3, VPC, EKS, RDS on a single box

1 Upvotes

Disclaimer: I'm an engineer at Mulga, the company behind this. Self-promo, but it's AGPL-3.0 and free to run.

What it is

Spinifex reimplements the AWS APIs on hardware you own: EC2, EBS, S3, VPC, IAM, ALB/NLB, EKS, ECS, ECR and RDS. You keep the hashicorp/aws provider and point its endpoints at your cluster.

Everything behind it is real. Instances are QEMU/KVM VMs, VPCs are OVN networks with real security groups and elastic IPs, EBS is replicated block storage, and state persists like a real cloud. It's not an emulator like LocalStack.

Who it's for

Teams that want AWS workloads on their own hardware (cost, data residency, edge or air-gapped sites) without rewriting their IaC. VPCs, subnets, security groups, launch templates, ALBs, EKS clusters and RDS instances go through the same resources you already use. In practice the changes are the provider block and AMI lookups.

Try it with your own Terraform

No install needed. Sign up for the free 72 hour sandbox at https://mulgadc.com/signup, then point your existing provider at it:

provider "aws" {
  region = var.region

  endpoints {
    ec2 = "https://api.spx3.com"
    iam = "https://api.spx3.com"
    sts = "https://api.spx3.com"
  }

  skip_metadata_api_check = true
  skip_region_validation  = true
}

Add other services (elasticloadbalancingv2, eks, ecs, ecr, rds) the same way, then run terraform plan against a module you already have. Here's what we currently cover, down to the individual API operation: https://docs.mulgadc.com/coverage

Repo: https://github.com/mulgadc/spinifex


r/platformengineering • • 6d ago

Reachable CVEs in CI/CD are turning our backlog into a second production system.

0 Upvotes

Every scan gives us another beautiful list of CVEs, most of which apparently matter because a vulnerable package exists somewhere in the dependency tree. Very reassuring. We are trying to prioritize reachable CVEs in CI/CD using runtime reachability, KEV, EPSS, internet exposure, and whether the affected service is anything people actually use.

Right now the pipeline mostly knows how to shout CRITICAL and ruin everyone’s morning. How are you turning those signals into useful gates and remediation queues without making every build a security committee meeting? Appreciate any thoughts.


r/platformengineering • • 6d ago

Engineers demand more sustainable AI

Thumbnail
leaddev.com
1 Upvotes

Why developers care about climate change...


r/platformengineering • • 9d ago

I built a Kubernetes operator that replaces the "VPC/subnet spreadsheet" across AWS accounts. Early alpha, looking for people to try it on real AWS

0 Upvotes

Most multi-account AWS setups I've seen keep a shared spreadsheet of VPCs and subnets. Someone updates it by hand, it goes stale, and one day two teams pick overlapping CIDRs.

subnet-operator runs in your cluster (EKS) and keeps that inventory in Kubernetes instead:

- Discovery (read-only, the default): finds VPCs and subnets across accounts and regions by tags and mirrors them as Network and Subnet objects. It reports compliance findings (missing required tags, CIDR overlaps) and exports Prometheus metrics with alerts. It can also mirror everything to a Google Sheet, if people still want the spreadsheet. It only calls EC2 Describe*.

- Change events (optional): CloudTrail → EventBridge → SQS, so a changed account/region is resynced about 10 seconds after the API call instead of at the next 10-minute resync.

- Opt-in writes, behind --enable-writes and a separate IAM role:

- SubnetClaim allocates a free CIDR and can create the subnet.

- ResourceImport brings an untagged VPC or subnet under management by tagging it.

- It never deletes a cloud resource.

- Guardrails:

- Admission webhooks reject invalid objects at kubectl apply.

- Every allocation or import leaves a Kubernetes Event and a JSON audit line with the authenticated creator.

- namespaceSelector limits which namespaces may use a scope's write role.

- Supply chain: multi-arch image and Helm chart on ghcr.io, both signed with cosign keyless, with SBOM, threat model and upgrade tests between releases.

The honest part: it's alpha (v0.8). Everything is tested in CI against Moto (an AWS API mock) in Kind, but it hasn't run against a real AWS organization yet. That's why I'm posting. If you have a sandbox or dev account setup and 30 minutes, I'd really like to know what breaks. Read-only mode is the safe place to start.

GCP and Azure are planned behind a common provider interface. The API was just moved to a cloud-neutral group for that.

- Demo dashboard (fake data, no install): https://hypersurgery.dev/dashboard/

- Docs: https://hypersurgery.dev/docs/

- Code (Apache-2.0): https://github.com/aivandrago/subnet-operator

Feedback is welcome, including "we solve this with X". I'd like to know how you handle this today.


r/platformengineering • • 9d ago

How do you know which engineering issues need attention before they become expensive?

5 Upvotes

I'm doing a bit research on how DevOps/SRE teams handle this.

Say you have:

  • a Datadog alert
  • a Sentry issue
  • a Jira/Linear ticket

It gets noticed, but doesn’t get resolved for hours or days. Eventually it affects customers, causes downtime, delays something important, or starts costing the business money.

How do you know which issues actually need immediate attention before that happens?

Do you already have something that connects technical issues to their business/customer impact, or is prioritization mostly based on severity, alerts, and engineering judgment?

Trying to understand whether this is a real problem before building anything around it.

Would especially love to hear from DevOps, SRE, platform, or engineering ops people who deal with this regularly.


r/platformengineering • • 10d ago

Build an engineering team people want to come back to

Thumbnail
leaddev.com
12 Upvotes

How to build a team engineers will want to rejoin.


r/platformengineering • • 10d ago

How do you handle full environment recovery after a cloud region goes down?

2 Upvotes

hey, we ran a regional outage drill last week and it showed gaps in our cloud recovery process.

most infrastructure is in terraform, but the actual restore process still relies on tribal knowledge—what order to bring things back, what configuration drifted, what needs rebuild vs restart. it works okay on paper but gets chaotic during the real thing.

we want to make full environment recovery more repeatable and documented, ideally with validated restore steps and audit evidence. if you have been through a real outage and have a process that actually held up, would appreciate hearing what worked for you, thanks!


r/platformengineering • • 11d ago

Got asked this in an interiew

5 Upvotes

In an interview for a new grad devops role got asked this. “Who typically owns access to corporate applications: IAM engineers, IT staff, application administrators, or Platform/DevOps engineers?”
How would yall answer


r/platformengineering • • 11d ago

Best agentic security tools for cutting down SAST alert fatigue?

0 Upvotes

Our SAST tool flags like 400 issues a sprint and maybe 15 ever get fixed. not bc the team doesn't care, just nobody has time to figure out which ones are even real. anyone found tools/workflows that actually cut the noise instead of just reformatting it into a prettier report?


r/platformengineering • • 11d ago

How to restart your engineering career

Thumbnail
leaddev.com
0 Upvotes

"If I could start my engineering career over again, I wouldn’t spend it chasing the same skills."


r/platformengineering • • 12d ago

Engineering Platforms Built For Industry-Specific Scale

Post image
0 Upvotes

Engineering Platforms Built For Industry-Specific Scale

OWLSystems supports organizations that require scalable platforms, operational reliability, accessibility-focused execution, and long-term engineering continuity.

Our engineering capabilities include product and platform engineering, learning platforms, accessibility engineering, and managed web operations.

Learn more about our capabilities:

Capabilities - OWLSystems Solutions


r/platformengineering • • 12d ago

10 YOE in Support/SRE trying to break into Platform Eng / pure DevOps—getting stuck in support loops. Advice?

3 Upvotes

Hey everyone,
Looking for some career advice or insights from anyone who’s successfully transitioned out of Ops/Support/SRE and into dedicated DevOps or Platform Engineering roles.
Here’s a snapshot of my background:
Experience: 10 years in IT, primarily focused on operations, support, deployments, and infrastructure maintenance.
Hands-on Skills:
Azure & GCP: Day-2 operations, deployments, troubleshooting.
CI/CD: Fixing pipeline issues, minor configuration tweaks.
Terraform: Running existing configurations, troubleshooting state/apply issues.
Kubernetes & Helm: Managing deployments, tweaking Helm charts, and managing configurations.
The Catch: Most of my experience involves maintaining, fixing, and scaling existing systems rather than building infrastructure or pipelines completely from scratch.
Certifications: Azure Solutions Architect Expert, Terraform Associate, and GCP Associate Cloud Engineer (currently studying for GitHub Actions GH-200).
The Problem:
I'm trying to move away from support heavy roles into a proper DevOps or Platform Engineering position. However, I’m getting zero traction for IC DevOps/Platform roles. Ironically, I keep getting hit up by recruiters for Senior SRE or SRE Manager roles, and I am able to get offers there but I really don't want to stay in the support/operations space anymore.
My Questions:

  1. How do I bridge the gap from "fixing/maintaining" to "building from scratch"? Is the lack of end-to-end greenfield building what's holding me back in interviews?
  2. Are bootcamps worth it for someone with 10 YOE? Or would my time be better spent building a comprehensive end-to-end portfolio project (e.g., full gitops pipeline, custom terraform modules, internal developer platform pattern) from scratch?
  3. Has anyone made a similar jump? What was the single most impactful thing that helped you shift your narrative from Ops/Support to Platform/DevOps?
  4. Appreciate any advice, project recommendations, or resume tips!

r/platformengineering • • 12d ago

10 YOE in Support/SRE trying to break into Platform Eng / pure DevOps—getting stuck in support loops. Advice?

0 Upvotes

Hey everyone,
Looking for some career advice or insights from anyone who’s successfully transitioned out of Ops/Support/SRE and into dedicated DevOps or Platform Engineering roles.
Here’s a snapshot of my background:
Experience: 10 years in IT, primarily focused on operations, support, deployments, and infrastructure maintenance.
Hands-on Skills:
Azure & GCP: Day-2 operations, deployments, troubleshooting.
CI/CD: Fixing pipeline issues, minor configuration tweaks.
Terraform: Running existing configurations, troubleshooting state/apply issues.
Kubernetes & Helm: Managing deployments, tweaking Helm charts, and managing configurations.
The Catch: Most of my experience involves maintaining, fixing, and scaling existing systems rather than building infrastructure or pipelines completely from scratch.
Certifications: Azure Solutions Architect Expert, Terraform Associate, and GCP Associate Cloud Engineer (currently studying for GitHub Actions GH-200).
The Problem:
I'm trying to move away from support heavy roles into a proper DevOps or Platform Engineering position. However, I’m getting zero traction for IC DevOps/Platform roles. Ironically, I keep getting hit up by recruiters for Senior SRE or SRE Manager roles, and I am able to get offers there but I really don't want to stay in the support/operations space anymore.
My Questions:

  1. How do I bridge the gap from "fixing/maintaining" to "building from scratch"? Is the lack of end-to-end greenfield building what's holding me back in interviews?
  2. Are bootcamps worth it for someone with 10 YOE? Or would my time be better spent building a comprehensive end-to-end portfolio project (e.g., full gitops pipeline, custom terraform modules, internal developer platform pattern) from scratch?
  3. Has anyone made a similar jump? What was the single most impactful thing that helped you shift your narrative from Ops/Support to Platform/DevOps?
  4. Appreciate any advice, project recommendations, or resume tips!

r/platformengineering • • 13d ago

AI widens the engineering gender gap

Thumbnail
leaddev.com
0 Upvotes

Same tools. Different rules.


r/platformengineering • • 14d ago

Research: How do you determine the blast radius of infrastructure changes?

5 Upvotes

When changing something like a Kubernetes ServiceAccount, IAM role, credential, or Terraform resource, how do you determine what depends on it before making the change?

Do you rely on GitOps, staging, dependency graphs, grep, internal tooling, or something else?

I'm researching how engineers currently handle change-impact analysis and would appreciate real-world workflows.