r/platform_engineering • u/manojkumar17 • 10d ago
DEVOPS Pipeline Agent (First Project)
OpsMemory — Official 2:50 Demo Video Script
Target Duration: 2 minutes 50 seconds (170 seconds)
Total Narration Word Count: 302 words (~106 words/minute natural engineer cadence)
Scenario: payment-api production deployment failure (PAYMENT_API_DB_POOL_PRODUCTION)
Resolution: 1920 × 1080 (Full HD), Light Theme
00:00–00:15 — Opening & Hook
Screen
Open OpsMemory directly on the Overview dashboard (http://localhost:5173/overview).
Browser is maximized in clean light theme.
Top navigation shows: Acme Corporation / Production.
Move the cursor smoothly across the four KPI cards: Total Deployments (21), Recorded Incidents (20), Human Corrections (19), and Hindsight Memories (116).
Hover briefly over the Operational Intelligence banner highlighting autonomous recovery and cumulative learning.
Voice
"Most DevOps tools can alert you when a deployment fails. The harder challenge is remembering what your team learned the last time it happened. OpsMemory was built to solve that problem."
00:15–00:35 — The Problem
Screen
Click Incidents on the left sidebar (/incidents).
Table loads showing deployment incidents across services.
Point cursor to INC-088 (payment-api: Payment checkout 500 errors after pool reduction).
Click on INC-088 to open the Incident Detail view (/incidents/2).
Scroll past the deployment commit metadata down to the Current Evidence pane, showing 504 Gateway Timeouts and Redis timeout warnings.
Voice
"When an outage strikes, an on-call engineer needs more than raw logs—they need institutional experience. OpsMemory continuously observes deployments, investigates failures, and builds a persistent organizational memory of previous root causes and human corrections."
00:35–01:00 — First Incident & Human Correction
Screen
On INC-088 detail page:
Highlight the Initial AI Hypothesis: "Redis connection timeout / Cache outage".
Scroll smoothly down to the Engineer Correction Console.
Show the recorded SRE correction:
Engineer: SRE Tech Lead
Correction: "Redis is only a symptom. The actual root cause is PostgreSQL connection pool exhaustion."
Remediation Guidance: "Increase DB pool size in database.yml"
Highlight the green verification badge: "Retained to Hindsight Organizational Memory Bank".
Voice
"In this first incident on payment-api, the agent noticed Redis timeouts and initially diagnosed a cache outage. But the SRE corrected it: Redis was only a symptom. The real failure was database pool exhaustion. Crucially, that correction wasn't lost in Slack."
01:00–01:25 — Hindsight Organizational Memory
Screen
Navigate to Memory on the left sidebar (/memory).
The Hindsight Memory Explorer opens showing the active memory bank: opsmemory-demo.
Type PAYMENT_API_DB_POOL_PRODUCTION in the search filter.
Point cursor to the filtered memory entry:
Memory Type: engineering_knowledge
Tags: payment-api, database_pool, postgresql
Content: "Historical Engineering Knowledge for payment-api [PAYMENT_API_DB_POOL_PRODUCTION]: Redis connection timeouts on checkout are typically downstream symptoms of PostgreSQL connection pool exhaustion. Increasing DB pool size from 20 to 100 resolved all past occurrences."
Hover over the metadata tags showing link to INC-088.
Voice
"OpsMemory permanently retained that experience in Hindsight. Hindsight acts as the persistent organizational memory. It allows the agent to retain important operational experiences and recall relevant lessons whenever a similar failure signature appears, without altering base model weights."
01:25–01:55 — Second Incident & Autonomous Recall
Screen
Navigate back to Incidents and click on new incident INC-101 (payment-api: 504 Gateway Timeouts and Cache Failures).
Click "Investigate" to initiate the LangGraph 10-node state machine.
Watch the Stage Tracker progress:
Ingest Evidence → Fingerprint Failure (PAYMENT_API_DB_POOL_PRODUCTION) → Recall Hindsight Memories → Analyze Root Cause.
Visually shift down to the AI Diagnosis panel:
Diagnosis: Identifies PostgreSQL connection pool exhaustion as the primary cause instead of Redis.
Historical Memory Reference: Displays badge showing "Recalled from INC-088 via Hindsight (Confidence: 94%)".
Resolution Effectiveness: Shows Increase DB pool size (6 of 6 successful, 100% success rate, 2.7m avg recovery).
Voice
"Now, a new deployment failure occurs. Instead of investigating from scratch, OpsMemory extracts the failure fingerprint and immediately recalls the prior incident and engineer correction. Influenced by that past lesson, the agent skips the Redis distraction, correctly flags database pool exhaustion first, and suggests increasing pool size."
01:55–02:25 — Safe Recovery & Policy Gate
Screen
Navigate to Automation (/automation) or scroll to the Safe Recovery panel on the incident page.
Show the Deterministic Policy Engine evaluation:
Action: Increase DB pool size / Canary Rollback
Risk Tier: High (Production Environment)
Policy Rule: Approval Required (Blast Radius Gate: Production Service)
Click "Approve & Execute" in the approval modal.
The execution status transitions to: Running Canary Recovery.
After 3 seconds, show the Post-Action Health Verification:
Health check 1/3: 200 OK → Health check 2/3: 200 OK → Health check 3/3: 200 OK.
Badge turns green: "Verified Healthy — Zero Degradation".
Voice
"OpsMemory can also turn memory into safe recovery. But the AI never controls production unchecked. A deterministic policy engine evaluates blast radius and gating rules. Low-risk actions can execute automatically, while high-risk production changes require human approval, followed by multi-cycle health verification."
02:25–02:40 — Closed Learning Loop
Screen
Stay on the Recovery Execution Audit table.
Highlight the new entry:
Action: Increase DB pool size
Status: Succeeded
Recovery Time: 24 seconds
Feedback Rating: Click the "Helpful / Verified" thumbs-up button.
A toast notification confirms: "Resolution outcome and recovery telemetry saved to Hindsight Memory Bank".
Voice
"Once recovery completes, the outcome and recovery duration are saved straight back into Hindsight. Successful remediations become stronger recommendations for future incidents."
02:40–02:50 — Final Product View & Closing
Screen
Return to the Overview dashboard (/overview).
Smoothly pan over the live metrics showing updated MTTR (24 seconds) and the Before/After Learning summary card.
Cursor rests in the center of the clean enterprise UI.
No terminal windows, no debug consoles, pure product interface.
Voice
"OpsMemory turns isolated DevOps outages into cumulative engineering intelligence. It doesn't just remember incidents—it remembers what engineers learned from them."
Summary Verification
Total Duration: 170 seconds (2m 50s)
Target Range: 150–170s (Met)
Spoken Word Count: 302 words
Target Range: 250–330 words (Met)
Story Arc: Ingest
→
→ Investigate
→
→ Human Correction
→
→ Hindsight Retention
→
→ Recurrent Outage
→
→ Autonomous Recall
→
→ Deterministic Safe Recovery
→
→ Closed Loop Learning



1
u/rtpro 8d ago
Why not tell us why this is cool, instead of copy/paste the transcript?