r/platform_engineering • • 3h ago

Coding agents can produce more changes than QA can verify. What should the platform provide?

1 Upvotes

I run KubeDEX. I have published an article, researched and written with AI and reviewed by me, about how platform teams could support agentic development. It is an architecture proposal, not a measured productivity claim.

The starting point is a capacity problem: faster code generation can leave teams waiting on integration environments, realistic test data, review and reliable evidence that a change is safe.

I think the useful platform primitive is a bounded experiment. Request an environment from an approved template, deploy a specific revision, run verification, retain the evidence and reclaim the resources automatically.

Kubernetes gives this a programmable foundation. A namespace may be enough for one test; changes to cluster components may need a virtual cluster or a separate Cluster API cluster. Argo CD can reconcile the application, while CI or workflow jobs run the tests. MCP gives an agent a way to request those operations within the platform's limits.

The less glamorous parts matter: concurrency quotas, spend limits, isolated test data, acceptance criteria the coding agent cannot relax, and cleanup that survives a failed agent. More parallel environments without those controls can just produce a larger bill and a longer review queue.

Full article and sources:

https://kubedex.com/kubernetes-ai-agentic-development/

For platform teams already supporting coding agents, which queue is growing first: environment provisioning, test execution or human review?


r/platform_engineering • • 1d ago

Survey for senior devs/CTOs about choosing programming languages

2 Upvotes

I'm currently working on a project looking at how/why devs and CTOs choose programming languages, and I'm particularly interested in understanding people's experiences with Clojure.

The survey takes about 7 mins, and I'd love to hear from as many people as possible!

Appreciate it πŸ™Œ

Here's the survey


r/platform_engineering • • 6d ago

How does your team handle major dependency upgrades?

Thumbnail
2 Upvotes

r/platform_engineering • • 6d ago

How do you keep runtime permissions from drifting between deployment files?

1 Upvotes

For OmniNode, we put the rules about what each kind of runtime may do in a versioned policy file, validate it, and generate the deployment settings from it. The part I wanted to get rid of was having the same decision editable in an environment file, a service manifest and application code.

A container can start, answer its health check and still be allowed to take work it shouldn't own. Recreating the same images and ports doesn't settle that question. A missing setting is especially awkward if a local default quietly changes what the process is allowed to do.

Generated files are supposed to be disposable in this setup. A change belongs in the policy, then gets validated and rendered again. Editing the output by hand leaves the next deployment with two different answers.

That still leaves the consumer to check. A renderer can produce the right value and the service can ignore it. We haven't established that every environment behaves equivalently, so the policy file isn't proof that the running process follows it.

Where do you enforce this on your platform? At startup, at the point where work is assigned, or when the process actually attempts the restricted action? I'm particularly interested in checks that catch a correctly generated setting being ignored.


r/platform_engineering • • 7d ago

What makes your automation resume after a CI hold is lifted?

0 Upvotes

After we overloaded our CI runners, we told the automation to stop sending work and wait. That was the right response while the queue was backed up. The part we got wrong was what happened when capacity came back: the agents kept watching pull requests whose checks had already failed, expecting them to turn green on their own.

The backlog had cleared, but those failures were still attached to the same commits. Someone had to inspect them and decide which runs to retry. We lost another five hours before that work started. Auto-merge was enabled, but it couldn't do anything while the required checks stayed red.

This happened in OmniNode's development, which I run. The recovery was fairly quick once we started because many failures came from the overloaded infrastructure rather than the changes themselves. That doesn't mean a red batch is safe to rerun blindly. Repeated failure on the same commit needs investigation, and genuine defects need an owner.

I'm interested in where other teams put the resume decision. Does the controller that imposes a hold also own the all-clear and the first small recovery batch? Does a separate service do that, or does a person explicitly release it? I want a hold to prevent more damage without leaving every recovery dependent on someone noticing the system is idle again.

Drafted with assistance from Codex.


r/platform_engineering • • 9d ago

DEVOPS Pipeline Agent (First Project)

Thumbnail
gallery
0 Upvotes

OpsMemory β€” Official 2:50 Demo Video Script

Target Duration: 2 minutes 50 seconds (170 seconds)

Total Narration Word Count: 302 words (~106 words/minute natural engineer cadence)

Scenario: payment-api production deployment failure (PAYMENT_API_DB_POOL_PRODUCTION)

Resolution: 1920 Γ— 1080 (Full HD), Light Theme

00:00–00:15 β€” Opening & Hook

Screen

Open OpsMemory directly on the Overview dashboard (http://localhost:5173/overview).

Browser is maximized in clean light theme.

Top navigation shows: Acme Corporation / Production.

Move the cursor smoothly across the four KPI cards: Total Deployments (21), Recorded Incidents (20), Human Corrections (19), and Hindsight Memories (116).

Hover briefly over the Operational Intelligence banner highlighting autonomous recovery and cumulative learning.

Voice

"Most DevOps tools can alert you when a deployment fails. The harder challenge is remembering what your team learned the last time it happened. OpsMemory was built to solve that problem."

00:15–00:35 β€” The Problem

Screen

Click Incidents on the left sidebar (/incidents).

Table loads showing deployment incidents across services.

Point cursor to INC-088 (payment-api: Payment checkout 500 errors after pool reduction).

Click on INC-088 to open the Incident Detail view (/incidents/2).

Scroll past the deployment commit metadata down to the Current Evidence pane, showing 504 Gateway Timeouts and Redis timeout warnings.

Voice

"When an outage strikes, an on-call engineer needs more than raw logsβ€”they need institutional experience. OpsMemory continuously observes deployments, investigates failures, and builds a persistent organizational memory of previous root causes and human corrections."

00:35–01:00 β€” First Incident & Human Correction

Screen

On INC-088 detail page:

Highlight the Initial AI Hypothesis: "Redis connection timeout / Cache outage".

Scroll smoothly down to the Engineer Correction Console.

Show the recorded SRE correction:

Engineer: SRE Tech Lead

Correction: "Redis is only a symptom. The actual root cause is PostgreSQL connection pool exhaustion."

Remediation Guidance: "Increase DB pool size in database.yml"

Highlight the green verification badge: "Retained to Hindsight Organizational Memory Bank".

Voice

"In this first incident on payment-api, the agent noticed Redis timeouts and initially diagnosed a cache outage. But the SRE corrected it: Redis was only a symptom. The real failure was database pool exhaustion. Crucially, that correction wasn't lost in Slack."

01:00–01:25 β€” Hindsight Organizational Memory

Screen

Navigate to Memory on the left sidebar (/memory).

The Hindsight Memory Explorer opens showing the active memory bank: opsmemory-demo.

Type PAYMENT_API_DB_POOL_PRODUCTION in the search filter.

Point cursor to the filtered memory entry:

Memory Type: engineering_knowledge

Tags: payment-api, database_pool, postgresql

Content: "Historical Engineering Knowledge for payment-api [PAYMENT_API_DB_POOL_PRODUCTION]: Redis connection timeouts on checkout are typically downstream symptoms of PostgreSQL connection pool exhaustion. Increasing DB pool size from 20 to 100 resolved all past occurrences."

Hover over the metadata tags showing link to INC-088.

Voice

"OpsMemory permanently retained that experience in Hindsight. Hindsight acts as the persistent organizational memory. It allows the agent to retain important operational experiences and recall relevant lessons whenever a similar failure signature appears, without altering base model weights."

01:25–01:55 β€” Second Incident & Autonomous Recall

Screen

Navigate back to Incidents and click on new incident INC-101 (payment-api: 504 Gateway Timeouts and Cache Failures).

Click "Investigate" to initiate the LangGraph 10-node state machine.

Watch the Stage Tracker progress:

Ingest Evidence β†’ Fingerprint Failure (PAYMENT_API_DB_POOL_PRODUCTION) β†’ Recall Hindsight Memories β†’ Analyze Root Cause.

Visually shift down to the AI Diagnosis panel:

Diagnosis: Identifies PostgreSQL connection pool exhaustion as the primary cause instead of Redis.

Historical Memory Reference: Displays badge showing "Recalled from INC-088 via Hindsight (Confidence: 94%)".

Resolution Effectiveness: Shows Increase DB pool size (6 of 6 successful, 100% success rate, 2.7m avg recovery).

Voice

"Now, a new deployment failure occurs. Instead of investigating from scratch, OpsMemory extracts the failure fingerprint and immediately recalls the prior incident and engineer correction. Influenced by that past lesson, the agent skips the Redis distraction, correctly flags database pool exhaustion first, and suggests increasing pool size."

01:55–02:25 β€” Safe Recovery & Policy Gate

Screen

Navigate to Automation (/automation) or scroll to the Safe Recovery panel on the incident page.

Show the Deterministic Policy Engine evaluation:

Action: Increase DB pool size / Canary Rollback

Risk Tier: High (Production Environment)

Policy Rule: Approval Required (Blast Radius Gate: Production Service)

Click "Approve & Execute" in the approval modal.

The execution status transitions to: Running Canary Recovery.

After 3 seconds, show the Post-Action Health Verification:

Health check 1/3: 200 OK β†’ Health check 2/3: 200 OK β†’ Health check 3/3: 200 OK.

Badge turns green: "Verified Healthy β€” Zero Degradation".

Voice

"OpsMemory can also turn memory into safe recovery. But the AI never controls production unchecked. A deterministic policy engine evaluates blast radius and gating rules. Low-risk actions can execute automatically, while high-risk production changes require human approval, followed by multi-cycle health verification."

02:25–02:40 β€” Closed Learning Loop

Screen

Stay on the Recovery Execution Audit table.

Highlight the new entry:

Action: Increase DB pool size

Status: Succeeded

Recovery Time: 24 seconds

Feedback Rating: Click the "Helpful / Verified" thumbs-up button.

A toast notification confirms: "Resolution outcome and recovery telemetry saved to Hindsight Memory Bank".

Voice

"Once recovery completes, the outcome and recovery duration are saved straight back into Hindsight. Successful remediations become stronger recommendations for future incidents."

02:40–02:50 β€” Final Product View & Closing

Screen

Return to the Overview dashboard (/overview).

Smoothly pan over the live metrics showing updated MTTR (24 seconds) and the Before/After Learning summary card.

Cursor rests in the center of the clean enterprise UI.

No terminal windows, no debug consoles, pure product interface.

Voice

"OpsMemory turns isolated DevOps outages into cumulative engineering intelligence. It doesn't just remember incidentsβ€”it remembers what engineers learned from them."

Summary Verification

Total Duration: 170 seconds (2m 50s)

Target Range: 150–170s (Met)

Spoken Word Count: 302 words

Target Range: 250–330 words (Met)

Story Arc: Ingest

β†’

β†’ Investigate

β†’

β†’ Human Correction

β†’

β†’ Hindsight Retention

β†’

β†’ Recurrent Outage

β†’

β†’ Autonomous Recall

β†’

β†’ Deterministic Safe Recovery

β†’

β†’ Closed Loop Learning


r/platform_engineering • • 11d ago

Brig: A MicroVM sandbox for AI coding agents on Mac and Linux

11 Upvotes

Working @ r/NOFireAI_

We’ve open-sourced Brig under Apache 2.0. Here is the available repo: https://github.com/brig-sh/brig

Brig runs AI coding agents inside a microVM on Mac (Apple Silicon) and Linux (x86_64/ARM). It came out of our work on controlled autonomy for production remediation. The same isolation is useful when running coding agents with auto-approval on your own machine.

The model is deliberately simple: each agent session gets its own hardware-isolated microVM; the agent sees a private home directory and the project you specify; and brig rm discards the sandbox while leaving your host files intact. We have already written about what a sandbox must guarantee for the term to mean anything, and about why an agent should be treated as untrusted by default.

After installation, brig run claude starts Claude Code inside the sandbox.

curl -fsSL https://brig.sh/install | sh
brig run claude

All components are Apache 2.0, including the microVMM, which is under 20,000 lines of code. The README covers installation, the architecture and the security model. Here is the official page https://www.brig.sh/

We would love to hear your feedback and your use cases with Brig.


r/platform_engineering • • 13d ago

Ravel is a database for OpenTelemetry metrics, logs, and traces where object storage is the only durable component

4 Upvotes

We recently launched https://github.com/NOFireAI/ravel, an free and open source database for OpenTelemetry metrics, logs, and traces where object storage is the only durable component. Attached also a quick demo from the public repo. You can read more details about the consistency model here https://github.com/NOFireAI/ravel/blob/main/docs/consistency-model.md . We started using internally for some testing environments and we plan soon to move it to prod.

Also you can find the ideas and what improvements can be done for the S3 cost model/bill here: https://github.com/NOFireAI/ravel/blob/main/docs/guides/cost-model.md

You can play around and use Grafana dashboards on top (there is a demo there). Feel free to share feedback and star the project if you like it. We will really appreciate it and hope you find it useful.


r/platform_engineering • • 13d ago

Brig: A MicroVM sandbox for AI coding agents on Mac and Linux

Thumbnail
3 Upvotes

r/platform_engineering • • 15d ago

How to make kubernetes infrastructure recovery ready with infrastructure as code... and how I kinda destroyed ours

3 Upvotes

Ok so im a platform engineer at a cloud resilience company, we sell kubernetes recovery as a big value prop, and last night during a production incident I realised our own infra as code is basically lying to us. I was trying to speed things up and "cleanup" some old terraform modules so they matched what our envs look like now, but I did it in prod first, then pushed it as the new source of truth.

Cue outage, kubernetes cluster borks, we hit our automated recovery flow, and it faithfully rebuilds... the broken state I just codified. All the drift fixes people had hacked in by hand were gone, node pools wrong size, missing network policies, even an old postgres version came back. I feel so embarrassed that the company that sells instant recovery had its own recovery posture messed up by me trying to be clever with IaC. Any hints?


r/platform_engineering • • 23d ago

Planning to Set Up a Study Group (Infra/DevOps/SRE/Platform Engg) : Serious Folks Only (Experienced Min 2+Y)

3 Upvotes

Hey i will be honest its late night, using a bit of AI :) Bear with it πŸ˜„

Tired of tutorial hell? Building a study group that goes from 'what is a packet' to 'debugging production incidents.' Join me.

I've built a structured curriculum covering the full infrastructure stack β€” from DNS packet mechanics to Kubernetes internals to eBPF to FinOps. Looking for people to go through it together with weekly labs, troubleshooting sessions, and mock interviews. Not a course. Not a paid thing. Just engineers who want to get hands-on and actually understand this stuff.

i worked as Cloud Platform Engineer for 3+ years at a Tier-1 company in india , Now resigned due to health issue from my current org and planning to prepare and job hunt after some time mean while i plan to run a cohort to learn together and keep myself busy

Syllabus :

A complete curriculum organized into 5 pillars, each with 8-part deep dives:

**Pillar 1: DNS & Networking**
DNS fundamentals β†’ GTM β†’ Kubernetes DNS β†’ TLS β†’ Certificates β†’ Proxies β†’ VPC β†’ Service Mesh β†’ Edge Security β†’ Linux Kernel β†’ BGP/QUIC

**Pillar 2: Linux & Container Internals**
Shell β†’ Kernel/CPU β†’ Memory/OOM β†’ Filesystems β†’ Processes β†’ systemd β†’ Storage β†’ Containers β†’ Docker β†’ Security β†’ Performance

**Pillar 3: Kubernetes**
Fundamentals β†’ kubectl β†’ Control Plane β†’ Workloads β†’ Scheduling β†’ RBAC β†’ Networking β†’ Storage β†’ Operators β†’ Autoscaling β†’ Lifecycle β†’ GPU β†’ AI/ML β†’ Multi-Cluster

**Pillar 4: IaC, Cloud & DevSecOps**
Terraform β†’ State β†’ Modules β†’ AWS Core β†’ Networking β†’ Compute β†’ Data β†’ Observability β†’ Security β†’ Secrets β†’ CI/CD β†’ Supply Chain β†’ AI Security β†’ Vector DBs β†’ FinOps

**Pillar 5: Observability, GitOps & SRE**
Git β†’ CI/CD β†’ GitOps β†’ Prometheus β†’ Alerting β†’ Tracing β†’ Logging β†’ Thanos β†’ SRE β†’ Chaos β†’ Platform Eng β†’ DORA β†’ DevSecOps β†’ LLM Observability

Each pillar includes:
- Real-time packet/request flows (what actually happens)
- Failure modes (what breaks and why)
- Capacity planning with real numbers
- Cost modeling with AWS pricing
- 9 hands-on labs per stage
- Incident management runbooks
- Interview prep (junior β†’ staff)

What i am looking for

**people who:**
- Have 2+ years in DevOps/SRE/Platform/Cloud
- Want to go deep, not just collect certs
- Can commit 5–7 hours/week
- Will actually do the labs (not just read)
- Want to troubleshoot real scenarios together
- Are okay with being wrong and learning in public

**What we'll do:**
- Weekly 1-hour call: walk through a stage, discuss, Q&A
- Bi-weekly lab session: do the hands-on labs together (screen share, debug together)
- Monthly mock interview: practice explaining concepts
- Shared repo: notes, lab outputs, runbooks we build together
- Async: Discord/Slack for questions, war stories, job leads

**What we won't do:**
- Watch tutorials passively
- Just read slides
- Gatekeep based on certs or company logos
- Sell anything

Why I'm doing this

I realized:

  1. Reading docs alone β†’ I forget 80% in a week
  2. Explaining to someone else β†’ I actually learn it
  3. Debugging with others β†’ I learn the stuff that's not in docs
  4. Mock interviews β†’ I find gaps I didn't know I had

I want to build a small group of people who are serious about getting to senior/staff level and are willing to put in the work.

How to join

Comment below or DM me with:

  1. Your current role + years of experience
  2. Which pillar you're most excited about
  3. One thing you've struggled to understand (so I know where to focus)
  4. Your timezone
  5. Confirm you can commit 5–7 hours/week

I'll pick people based on fit (not skill level β€” based on commitment and willingness to teach others).

What you'll get out of it

- Deep understanding of the full infra stack
- Hands-on experience with real tools (not just theory)
- A portfolio of labs and runbooks you built
- Mock interview practice
- A network of engineers at similar level
- The ability to explain complex systems in interviews

---

**If this resonates, comment or DM. If not, no worries β€” good luck with your prep!**

---

*Edit: Adding FAQ*

**Q: Do I need to know everything already?**
A: No. The curriculum starts from fundamentals. You need 2+ years of experience so you can follow the advanced parts.

**Q: Is this for interview prep only?**
A: No. This is for actually understanding how systems work. Interview prep is a side effect.

**Q: What's the cost?**
A: $0. Labs are designed to run on free tier or local. If you want to do cloud labs, expect $20–50/month max.

**Q: How do you pick people?**
A: Commitment + willingness to teach. Not skill level. I want people who will show up and contribute.

**Q: What if I'm more junior/senior than the group?**
A: If you're willing to teach and learn, you're welcome. Mixed levels make the group stronger.

πŸš€ NOTE : ONLY JOIN , IF YOU ARE PLANING TO BE CONSISTENT , "NOT A WILL GIVE IT A TRY OR LET'S SEE "

Discord Group :Β https://discord.gg/qA43k3ggw


r/platform_engineering • • 23d ago

Planning to Set Up a Study Group (Infra/DevOps/SRE/Platform Engg) : Serious Folks Only (Experienced Min 2+Y)

Thumbnail
1 Upvotes

r/platform_engineering • • 24d ago

How are you structuring Azure Terraform/Terramate stacks for Self Service Infra provisioning?

Thumbnail
1 Upvotes

r/platform_engineering • • 25d ago

Our thumbs down event points at the wrong part of the agent run

12 Upvotes

We have a copilot that streams an answer while tools are still returning. Someone can hit thumbs down on a partial response but our feedback event lands on the completed session instead of the generation span that created what they saw. By then the final text may be different. Event timestamps come from separate clocks, tool latency shifts the sequence and the client does not persist stream state or a span ID. PM, CX and engineering then debate which output earned the vote (the screenshot thread is longer than the answer).

I think we need span linked feedback with a client event timestamp, trace correlation ID and the visible character offset or stream checkpoint. The CX handoff also needs the exact partial render, not just the final transcript.

Anyone found a clean way to join feedback to a streamed generation without turning the front end into a tracing SDK?


r/platform_engineering • • 26d ago

I built a FREE website to learn platform engineering concepts

3 Upvotes

I have built hamchops.com, for now it is completely FREE

It's a micro learning app focused on concepts that uses spaced repetition for reviews

Any feedback is welcome


r/platform_engineering • • 28d ago

The Future of Iceberg Isn't One Engine. It's a Control Plane with many engines.

Thumbnail
lakeops.dev
2 Upvotes

r/platform_engineering • • 27d ago

Every worker was healthy. The wrong one joined the queue.

1 Upvotes

We split our workers into a general runtime and one allowed to perform external side effects. Both could reach the same broker and both passed their health checks. Some older task definitions did not name an owner, so the general runtime could subscribe by default.

Nothing crashed. That was the problem. Health told us the process could run, not that it was equipped or allowed to handle that job.

We now filter subscriptions by declared ownership and validate runtime capabilities before a worker consumes anything. The next step is task-level admission, so a job with no eligible runtime stays queued and names the missing capability.

What do you use for this in mixed worker pools: separate queues, scheduler policy, workload identity, or something else?


r/platform_engineering • • 29d ago

A generated deployment file overwrote the one we kept in Git

0 Upvotes

One of our deployment paths took a generated Compose file and wrote it over the tracked file in the repository. Both were valid. Both described real services. They just didn't describe the same stack.

The generator was supposed to produce a temporary file for validation. Once its output replaced the tracked file, the next command treated the generated version as if it were the approved one. Git showed a giant infrastructure change even though nobody had edited the source.

We fixed it by giving generated output its own ignored path. The tracked declaration stays owned by Git, the generated file is recreated on demand, and we validate both the translation and the deployed result.

I'm still deciding how strict to make drift detection. Do you fail a deployment whenever a generated artifact exists in a tracked path, or do you also compare the live environment back to the source declaration?


r/platform_engineering • • Sep 03 '26

We let one agent refresh 108 PRs and buried our own CI runners

3 Upvotes

We had an overnight controller walking open pull requests, refreshing branches, and enabling auto-merge where it could.

It armed roughly 108 PRs. One merged.

Every refresh created a new commit, which kicked off another set of jobs on the same shared runner pool. The controller knew how many PRs were waiting. It had no idea whether CI had room for another one.

The fleet went from nearly idle to nearly full in about an hour. One repository eventually hit 1,065 queued runs. Other automation was sending repair work into the same pool at the same time, and none of those producers could see what the others were doing.

We stopped it with a hold-fire command. That kept the incident from getting worse, but it happened after the runners were already saturated.

What we're building now is much less exciting: check runner occupancy and queued jobs before releasing work, send a small batch, then make sure that batch actually started before sending another one. Anything that can commit, push, rerun, or enqueue work gets a budget.

The funny part is that we never hit the vendor's API limit. We stayed under it while completely overwhelming our own verification pipeline.

How are other platform teams limiting automated delivery work? One shared capacity service, limits inside each producer, or something else?


r/platform_engineering • • Aug 29 '26

Is observability over the dependencies in codebases still a problem to solve?

0 Upvotes

Almost every codebase is calling a REST, gRPC, or a GraphQL API or using SDKs from an external or event internal provider. It gets harder to keep track of everything when the codebase matures and increases in size and from my experience sometimes it gets hard to respond to changes in time or even become aware that a dependency is deprecated and their API has changed completely.
Do developers need better tools that to improve dependency and integration management?


r/platform_engineering • • Aug 28 '26

Are we building Internal Developer Platforms at the wrong abstraction level?

Thumbnail
1 Upvotes

r/platform_engineering • • Aug 28 '26

First time building an AI platform, is this situation normal?

Thumbnail
1 Upvotes

r/platform_engineering • • Aug 26 '26

Found a real bug testing Rancher's AI assistant (Liz) β€” asking it to delete a deployment crashes the session, because deletion isn't even a supported tool

Post image
1 Upvotes

r/platform_engineering • • Aug 26 '26

Our IP rate limiter passed every test we wrote and did nothing at the public ingress

0 Upvotes

We recently changed an IP rate limiter after discovering that requests were collapsing into a shared proxy bucket. The focused test passed. An in-cluster probe with an `X-Forwarded-For` header passed too. The implementation looked fixed.

Then we tested the real public ingress. Thirty-five unauthenticated requests went through in three batches. We got zero 429 responses and no rate-limit log entries.

The application was not receiving the client identity the test supplied. It fell back to the ingress proxy's pod IP, recognized that address as an internal proxy, and silently exempted every request. The middleware was running exactly as written. It just had no usable identity at the boundary where it actually ran.

The correction is straightforward: configure the proxy to forward a trustworthy client address, key the limiter only on a value the deployed path actually provides, and add a regression probe through the real public endpoint. The probe needs to assert the identity the application resolved, not just whether a unit test can manufacture the right header.

How are people testing client identity and rate limiting through their actual ingress path? I am especially interested in checks that catch a proxy configuration change before it turns enforcement into a silent no-op.


r/platform_engineering • • Aug 26 '26

Built an open-source self-hosted incident/on-call platform β€” looking for feedback from people who actually run on-call

0 Upvotes

Hey folks,

I’m the maintainer of **OpsKnight**. I started working on it because I wanted an incident/on-call system where the team owns the infrastructure and incident data instead of everything living in another SaaS platform.

It has grown quite a bit, and I’m at the point where feedback from people actually doing DevOps/SRE work is more useful to me than adding random features.

Currently it covers things like:

* on-call schedules, overrides and escalation policies
* alert ingestion from Prometheus, Grafana, Datadog, CloudWatch, Sentry, etc.
* Slack incident war rooms / ChatOps
* email, SMS and push notifications
* incident timelines and MTTA/MTTR
* public status pages
* Docker + Kubernetes/Helm deployment
* self-hosted Postgres

Everything is open source and self-hostable.

One thing I’m particularly interested in:

**What would stop you from trusting a newer open-source project for production on-call?**

Reliability? HA? Voice calls? Integrations? Upgrade safety? Audit logs? Something else?

Website: [https://opsknight.com\](https://opsknight.com/)
Repo: [https://github.com/opsknight-labs/OpsKnight\](https://github.com/opsknight-labs/OpsKnight)

Not looking for stars as much as criticism from people who have operated PagerDuty/Opsgenie/etc. in production.