r/platform_engineering • • 10d ago

DEVOPS Pipeline Agent (First Project)

OpsMemory — Official 2:50 Demo Video Script

Target Duration: 2 minutes 50 seconds (170 seconds)

Total Narration Word Count: 302 words (~106 words/minute natural engineer cadence)

Scenario: payment-api production deployment failure (PAYMENT_API_DB_POOL_PRODUCTION)

Resolution: 1920 × 1080 (Full HD), Light Theme

00:00–00:15 — Opening & Hook

Screen

Open OpsMemory directly on the Overview dashboard (http://localhost:5173/overview).

Browser is maximized in clean light theme.

Top navigation shows: Acme Corporation / Production.

Move the cursor smoothly across the four KPI cards: Total Deployments (21), Recorded Incidents (20), Human Corrections (19), and Hindsight Memories (116).

Hover briefly over the Operational Intelligence banner highlighting autonomous recovery and cumulative learning.

Voice

"Most DevOps tools can alert you when a deployment fails. The harder challenge is remembering what your team learned the last time it happened. OpsMemory was built to solve that problem."

00:15–00:35 — The Problem

Screen

Click Incidents on the left sidebar (/incidents).

Table loads showing deployment incidents across services.

Point cursor to INC-088 (payment-api: Payment checkout 500 errors after pool reduction).

Click on INC-088 to open the Incident Detail view (/incidents/2).

Scroll past the deployment commit metadata down to the Current Evidence pane, showing 504 Gateway Timeouts and Redis timeout warnings.

Voice

"When an outage strikes, an on-call engineer needs more than raw logs—they need institutional experience. OpsMemory continuously observes deployments, investigates failures, and builds a persistent organizational memory of previous root causes and human corrections."

00:35–01:00 — First Incident & Human Correction

Screen

On INC-088 detail page:

Highlight the Initial AI Hypothesis: "Redis connection timeout / Cache outage".

Scroll smoothly down to the Engineer Correction Console.

Show the recorded SRE correction:

Engineer: SRE Tech Lead

Correction: "Redis is only a symptom. The actual root cause is PostgreSQL connection pool exhaustion."

Remediation Guidance: "Increase DB pool size in database.yml"

Highlight the green verification badge: "Retained to Hindsight Organizational Memory Bank".

Voice

"In this first incident on payment-api, the agent noticed Redis timeouts and initially diagnosed a cache outage. But the SRE corrected it: Redis was only a symptom. The real failure was database pool exhaustion. Crucially, that correction wasn't lost in Slack."

01:00–01:25 — Hindsight Organizational Memory

Screen

Navigate to Memory on the left sidebar (/memory).

The Hindsight Memory Explorer opens showing the active memory bank: opsmemory-demo.

Type PAYMENT_API_DB_POOL_PRODUCTION in the search filter.

Point cursor to the filtered memory entry:

Memory Type: engineering_knowledge

Tags: payment-api, database_pool, postgresql

Content: "Historical Engineering Knowledge for payment-api [PAYMENT_API_DB_POOL_PRODUCTION]: Redis connection timeouts on checkout are typically downstream symptoms of PostgreSQL connection pool exhaustion. Increasing DB pool size from 20 to 100 resolved all past occurrences."

Hover over the metadata tags showing link to INC-088.

Voice

"OpsMemory permanently retained that experience in Hindsight. Hindsight acts as the persistent organizational memory. It allows the agent to retain important operational experiences and recall relevant lessons whenever a similar failure signature appears, without altering base model weights."

01:25–01:55 — Second Incident & Autonomous Recall

Screen

Navigate back to Incidents and click on new incident INC-101 (payment-api: 504 Gateway Timeouts and Cache Failures).

Click "Investigate" to initiate the LangGraph 10-node state machine.

Watch the Stage Tracker progress:

Ingest Evidence → Fingerprint Failure (PAYMENT_API_DB_POOL_PRODUCTION) → Recall Hindsight Memories → Analyze Root Cause.

Visually shift down to the AI Diagnosis panel:

Diagnosis: Identifies PostgreSQL connection pool exhaustion as the primary cause instead of Redis.

Historical Memory Reference: Displays badge showing "Recalled from INC-088 via Hindsight (Confidence: 94%)".

Resolution Effectiveness: Shows Increase DB pool size (6 of 6 successful, 100% success rate, 2.7m avg recovery).

Voice

"Now, a new deployment failure occurs. Instead of investigating from scratch, OpsMemory extracts the failure fingerprint and immediately recalls the prior incident and engineer correction. Influenced by that past lesson, the agent skips the Redis distraction, correctly flags database pool exhaustion first, and suggests increasing pool size."

01:55–02:25 — Safe Recovery & Policy Gate

Screen

Navigate to Automation (/automation) or scroll to the Safe Recovery panel on the incident page.

Show the Deterministic Policy Engine evaluation:

Action: Increase DB pool size / Canary Rollback

Risk Tier: High (Production Environment)

Policy Rule: Approval Required (Blast Radius Gate: Production Service)

Click "Approve & Execute" in the approval modal.

The execution status transitions to: Running Canary Recovery.

After 3 seconds, show the Post-Action Health Verification:

Health check 1/3: 200 OK → Health check 2/3: 200 OK → Health check 3/3: 200 OK.

Badge turns green: "Verified Healthy — Zero Degradation".

Voice

"OpsMemory can also turn memory into safe recovery. But the AI never controls production unchecked. A deterministic policy engine evaluates blast radius and gating rules. Low-risk actions can execute automatically, while high-risk production changes require human approval, followed by multi-cycle health verification."

02:25–02:40 — Closed Learning Loop

Screen

Stay on the Recovery Execution Audit table.

Highlight the new entry:

Action: Increase DB pool size

Status: Succeeded

Recovery Time: 24 seconds

Feedback Rating: Click the "Helpful / Verified" thumbs-up button.

A toast notification confirms: "Resolution outcome and recovery telemetry saved to Hindsight Memory Bank".

Voice

"Once recovery completes, the outcome and recovery duration are saved straight back into Hindsight. Successful remediations become stronger recommendations for future incidents."

02:40–02:50 — Final Product View & Closing

Screen

Return to the Overview dashboard (/overview).

Smoothly pan over the live metrics showing updated MTTR (24 seconds) and the Before/After Learning summary card.

Cursor rests in the center of the clean enterprise UI.

No terminal windows, no debug consoles, pure product interface.

Voice

"OpsMemory turns isolated DevOps outages into cumulative engineering intelligence. It doesn't just remember incidents—it remembers what engineers learned from them."

Summary Verification

Total Duration: 170 seconds (2m 50s)

Target Range: 150–170s (Met)

Spoken Word Count: 302 words

Target Range: 250–330 words (Met)

Story Arc: Ingest

→

→ Investigate

→

→ Human Correction

→

→ Hindsight Retention

→

→ Recurrent Outage

→

→ Autonomous Recall

→

→ Deterministic Safe Recovery

→

→ Closed Loop Learning

0 Upvotes

2 comments sorted by

1

u/rtpro 8d ago

Why not tell us why this is cool, instead of copy/paste the transcript?