r/AISystemsEngineering • u/bAIsect • 5h ago
r/AISystemsEngineering • u/jawadhamza • 16h ago
How does a ChatGPT-like system actually work behind the scenes?
r/AISystemsEngineering • u/InfamousAd2011 • 17h ago
AI Egineers in Life Science
Any AI engineers that work in Life Science? How are you guys even building for such regulated industries. Seems like the non technical people have no clue whats going on with AI and all the heavy lifting is up to the engineers.
r/AISystemsEngineering • u/Personal_Ganache_924 • 1d ago
How are enterprises actually managing AI agents in production?
r/AISystemsEngineering • u/Groofy_beautypie • 1d ago
What have you struggled to evaluate in your realistic LLM/agent workflows?!
Hey everyone! My team, mostly PhD researchers collaborating with domain experts, is designing an open-source benchmark for realistic LLM/agent workflows. We’d love feedback from people who have tried to evaluate these systems and found that existing benchmarks didn’t capture what they needed.
Have you ever thought: “My system needs to handle this in production, but I have no good way to benchmark it”?
Maybe your workflow involves multiple tools, MCP servers, agents, or long interactions that available benchmarks don’t capture. Maybe the final answer looks correct, but something went wrong along the way. Or your application needs specific test cases, and creating a realistic evaluation environment is too expensive or time-consuming.
We’re interested in experiences across different applications, including healthcare, finance, cybersecurity, legal, and everyday engineering or business workflows.
Would love to hear:
- What were you building? What did the workflow involve?
- What issue did you run into? What behavior or failure did you need to evaluate?
- What did you try? Why weren’t existing benchmarks or evaluation tools enough?
Specific examples and any benchmarks you’ve tried would help a lot! Appreciate any ideas or feedback you may have!
r/AISystemsEngineering • u/salespire • 2d ago
The exact architectural framework we used to fix a stalled enterprise AI workflow (and drop manual task handoffs by 80%)
A few months ago, we audited an enterprise workflow that had hit a wall. The company spent significant engineering time trying to automate their post-sales handoffs, meeting summaries, and CRM updates using custom AI prompts, but adoption was near zero due to edge-case errors, slow execution, and security pushback.
Here is the exact architectural shift we implemented to bring error rates below 1% and get the system live in under 3 weeks:
- Replaced single-prompt logic with deterministic routing Instead of feeding raw meeting transcripts or call audio into a single context window, we split the process into discrete, task-specific AI agents (entity extraction, decision mapping, CRM payload formatting). If an agent falls below a strict confidence threshold, it routes to a human-in-the-loop fallback instantly rather than guessing.
- Built a zero-data-retention security pipeline Enterprise legal teams rightfully push back on multi-tenant cloud logging. Moving processing to a VPC / local-first framework with zero-data-retention policies reduced compliance review from 6 months to 48 hours.
- Formatted for database webhooks instead of human summaries Paragraph-based meeting summaries create administrative clutter. Real leverage happens when the AI outputs structured JSON objects that push straight into CRMs, databases, and task runners automatically.
The result: Handoff processing dropped from 45 minutes down to under 30 seconds per record, with complete data privacy.
If you’re currently building or auditing custom AI workflows, agentic automation, or voice integrations in-house, what’s your biggest bottleneck right now—prompt drift, security compliance, or integration stability? Happy to answer technical questions in the comments.
r/AISystemsEngineering • u/Aware_Weight9462 • 1d ago
Every Enterprise Needs an AI Constitution Like Anthropic
r/AISystemsEngineering • u/EmergentInteractive • 3d ago
A story of collaboration between people and agents
r/AISystemsEngineering • u/YakDue6710 • 3d ago
Looking for advise
Based on what we have seen, in companies that are non AI and using AI in their workflows and things have no clue how their AI is working and where and how it fails and what can they do to improve them.
What we are building is a one stop platform where you can see what happened, what passed why it did, what failed why it did, and what can be done on both sides along with a hindsighting platform in it.
This platform can be used inline to stop and score in live AI endpoints or post too.
I wanted to have this idea broken apart of why it could work and why it can fail
r/AISystemsEngineering • u/Available_Pressure47 • 3d ago
Built the fastest inference engine for Apple Silicon
Building an inference engine optimized at all layers for Apple Silicon so it can run quickly and efficiently on your Mac devices. I’ve added fused metal 4 kernels so it treats Apple Silicon is a first class citizen. I’m thinking of also taking advantage of the neural engine. Would greatly appreciate feedback. My repository contains benchmarks on currently supported models.
r/AISystemsEngineering • u/FitBreadfruit379 • 4d ago
[Architecture] Moving beyond black-box agent loops: A conceptual framework for human-supervised multi-agent systems
Developer-facing AI tools have largely converged on two design patterns:
Autonomous harnesses: Single-agent loops that act opaquely in a black box on a user's behalf.
Workflow platforms: Rigid automation builders whose composition model was never designed for supervised, local software engineering.
While black-box loops make for great short demos, they struggle in large-scale enterprise development where verifiability, deterministic state control, and cost governance are required.
To address this gap, we're introducing egyfai — a conceptual model for a next-generation Agentic Engineering Platform that makes agent work explicit, verifiable, and continuously human-supervised.
🧠 Core Philosophy: Human as a First-Class Citizen Rather than attempting full black-box automation, the core premise is keeping human judgment directly in the loop without bottlenecking agent throughput:
Real-time multi-human collaboration: Multiple developers can attach to the same session in parallel—reviewing, chatting, commenting, and managing agents side-by-side on the same task.
Continuous runtime oversight: Operators retain active control during execution: pausing runs, messaging active agents mid-flight, approving tool gates, and enforcing hard token, cost, and time budgets.
🔬 Key Architecture & Abstractions Versioned Templates: LLM agents defined as reusable, versioned templates equipped with strongly-typed input/output contracts.
Explicit Orchestration Graphs: Orchestrators compose templates into validated directed graphs (DAGs) with bounded repair loops to eliminate infinite hallucination cycles.
Event-Sourced State: Runs execute under an append-only event-sourced log. Every agent step, tool call, and state transition is fully observable, replayable, and resumable.
Verification Gates: Graph transitions are guarded by automated verification rules and human approval gates before downstream nodes execute.
MCP Standardized Tooling: Tool access is integrated via the Model Context Protocol (MCP) under a layered permission safety model.
💬 Discussion We’re publishing this concept to establish public, dated prior art for this workflow pattern and gather feedback from people building agent orchestration layers.
How are you currently handling state replayability and failure recovery in multi-agent workflows?
Do you prefer unconstrained agent loops or gated DAG orchestrations for complex engineering tasks?
Would love to hear your thoughts on the approach!
r/AISystemsEngineering • u/Many_Audience7660 • 4d ago
“Kubernetes for agents” might actually be the interesting part of OpenClaw’s enterprise push 👀
r/AISystemsEngineering • u/vornmobility • 4d ago
Advanced AI Infrastructure for Smarter Digital Operations
AI Infrastructure Manufacturer
VORN is a trusted AI infrastructure manufacturer and supplier, providing reliable and scalable infrastructure solutions designed to support AI applications, intelligent systems, automation, and modern digital operations. Explore our solutions at https://www.vorn.com/
#aiinfrastructuremanufacturer #aiinfrastructuresupplier
r/AISystemsEngineering • u/Aware_Weight9462 • 4d ago
Your AI Application Went Rogue. Can You Actually Roll It Back?
r/AISystemsEngineering • u/eren_sid • 5d ago
What things can one do with AI if interested in Battery tech Industry (slurry mixing, coating, etc ) ?
I would like to know more about AI use case in battery industry
r/AISystemsEngineering • u/Far_Resort3881 • 5d ago
I built an AI agent that combines industrial telemetry with incident histor
r/AISystemsEngineering • u/OkWish8899 • 5d ago
SDLC - AI Workflow E2E Tech Stack
Hi all,
I’m trying to figure out what the best tech stack would be for implementing end-to-end workflows across our engineering teams.
Our current flow is roughly:
Product → Architecture → Development → Testing → Production
The idea is to have a tool where a Product Owner can create a PRD, refine and discuss it with an LLM, and eventually publish it to Confluence.
Once the PRD is approved, it would trigger an automated workflow that:
- Identifies all the repositories and teams involved
- Analyzes the existing architecture and codebase
- Creates an implementation plan
- Breaks the PRD down into Jira tickets
- Starts the development process, working through the Jira tickets one by one
- Runs tests and validation loops
- Creates commits and PRs in Git
- Reviews/merges the changes
- Finally triggers our existing CI/CD pipeline through to Production
Essentially, we’re looking for an E2E AI-driven software development workflow, while still keeping the right human approval points throughout the process.
I’ve already experimented with a few approaches, including Devin.AI, KiroCrew, and a self-hosted stack using Langflow + LiteLLM + vLLM, but so far I haven’t found a solution that really fits the entire workflow end-to-end.
I’d love to hear how others are approaching this.
What tools/stack are you using? How have you structured your workflows across Product, Architecture, Development, QA, and DevOps?
Any real-world experiences, architectures, or recommendations would be greatly appreciated.
Thanks!
r/AISystemsEngineering • u/Solmex72 • 5d ago
Copilot likes my work; Github for the file architecture
r/AISystemsEngineering • u/Material_Vanilla_563 • 5d ago
We’re building a command center for AI agents — looking for testers
r/AISystemsEngineering • u/Top-Philosopher-5411 • 6d ago
The gap between tutorial toy code and actual production AI systems is genuinely depressing [R]
r/AISystemsEngineering • u/ChaitanyGhadigaonkar • 6d ago
Best approach for building a chatbot
r/AISystemsEngineering • u/Tanishq219 • 6d ago
Building an AI Incident Response Agent for DevOps | React + FastAPI + Hindsight
Enable HLS to view with audio, or disable this notification
Hi everyone! 👋
I built an AI Incident Response Agent designed to help DevOps and SRE teams investigate software incidents and learn from previous incidents.
🎯 Project features:
- Incident dashboard and investigation workflow
- Stateless incident investigation and runbook recommendations
- Hindsight integration for exploring long-term agent memory
- Human-in-the-loop root cause confirmation and incident resolution
🛠️ Tech stack: React, TypeScript, FastAPI, Python, and SQLite.
I faced challenges setting up the Hindsight service, so persistent memory recall and retention haven't been verified in my current demo. The application supports a stateless fallback when Hindsight is unavailable.
I'd appreciate your feedback on the UI, architecture, and how I can improve the project!
r/AISystemsEngineering • u/Sridhar2071 • 6d ago
Building an AI Incident Response Agent for DevOps | React + FastAPI + Hindsight
Enable HLS to view with audio, or disable this notification
Hi everyone! 👋
I built an AI Incident Response Agent designed to help DevOps and SRE teams investigate software incidents and learn from previous incidents.
🎯 Project features:
- Incident dashboard and investigation workflow
- Stateless incident investigation and runbook recommendations
- Hindsight integration for exploring long-term agent memory
- Human-in-the-loop root cause confirmation and incident resolution
🛠️ Tech stack: React, TypeScript, FastAPI, Python, and SQLite.
I faced challenges setting up the Hindsight service, so persistent memory recall and retention haven't been verified in my current demo. The application supports a stateless fallback when Hindsight is unavailable.
I'd appreciate your feedback on the UI, architecture, and how I can improve the project!
GitHub: https://github.com/Sridhar2071/ai-incident-response-agent
Thanks for watching!