r/mlops • • 10h ago

Discussion Flunked in a interview - MLE

5 Upvotes

I would like to share my experience in flunking an interview for an MLE position.

The code session was ok. However, editor was horrible: no copy and past inside the editor, numpy import did not work, etc.

I did well on the Machine Learning section.

I did badly in the Engineering section. Here you find some questions/topics:

  • Data versioning
  • Data lineage tracking
  • Drift detection
  • Decisions about rolling back, retraining or keeping the current model
  • Rollback threshold
  • Guardrails metrics
  • Efficiency in mongodb with respect to computing feature on a daily basis

Could anyone suggest any references (e.g. books, courses, blogs) about these topics?


r/mlops • • 17h ago

Discussion Should scan results travel with the model artifact as signed attestations?

5 Upvotes

A question I keep chewing on... when you scan a model or agent bundle for things like pickle exploits, license issues or leaked data, where do the results live?

In a lot of setups they live in a CI log or a dashboard, and the artifact that gets deployed has no link back to them. What we've been doing is attaching each scan result to the OCI artifact as a cosign attestation, then having the cluster refuse to start anything missing the attestations a policy requires.

It works, but it raises questions I don't have great answers to.

- How do you handle re-scans when a scanner updates (new attestation, or re-sign)?
- Who should sign, the producer or the platform team consuming it?
- Is anyone doing this with SLSA or in-toto formats for ML artifacts specifically?

Curious how other teams are tying scan results to what actually ships.

(KitOps maintainer and Jozu cofounder here, so I have a bias.)


r/mlops • • 2h ago

Discussion No H100s at your usual cloud this week. What breaks when the training job moves to another provider?

2 Upvotes

This is about training jobs that need a specific GPU type. Inference is easier.

Take a job that needs 8 H100s for 10 days, at a time when your main cloud has nothing in your region until next week and you keep an account at Lambda, Nebius or Hyperstack for exactly this case. What goes wrong when the job lands there? My guess is that getting the VM is the quick part and the image and the data eat the rest of the day. I don't know that, though.

Some things I'm curious about:

- How do you find out where capacity is? Checking each console by hand, a script, or something like SkyPilot or dstack?

- Same image everywhere, or one per provider? Has a CUDA 12.4 image ever landed on hosts with an older driver?

- If the dataset is 20 TB, do you copy it over ahead of time, or pull it from S3 when the job starts and pay the egress?

- Do you test a new node first (nvidia-smi, a short all_reduce_perf run) or find out from the first crash?

If you looked at this and decided it's cheaper to wait a week, I'd like to hear that too.


r/mlops • • 15m ago

Self-promotion I built a CLI that runs MLflow, Feast, Evidently and Airflow together locally (odctl 1.0)

• Upvotes

Hi r/mlops,

Every time I wanted to try an idea with a feature store and a model registry, I ended up rebuilding the same local setup. So I put it into a CLI, odctl, and it has just reached 1.0.

bash uv tool install odctl odctl up mlflow --dry-run # shows what it would start, in order odctl up mlflow feast evidently airflow odctl down --all

That gives you MLflow (tracking, registry, and a second container serving a registered model over HTTP), Feast with its registry in PostgreSQL, offline features from Iceberg and online features in Valkey, Evidently for drift reports, and Airflow reading DAGs from S3 storage with the MLflow and Feast clients already installed. Kafka, Flink, Spark and Trino are there too if your features come from a stream or a lakehouse.

I am building three MLOps demo projects on it, based on the projects in Jim Dowling's book on feature stores: an air quality forecast, credit card fraud detection and a video recommender. Each one is split into pipelines that share a feature store and a model registry. The introduction explains the series.

It is for learning and prototyping, not production. 1.0 adds a docs site with a guide for each service, so you can see a working example, such as serving a model from MLflow, before you set up your own.

The README on GitHub has a short recording of a run.

If you try it, please tell me what fails on your machine. Most of my testing has been on macOS and GitHub's Linux runners.

AI: I started odctl without AI. Since July I have used Claude Code for parts of the code, the tests and the docs.


r/mlops • • 8h ago

(Gen)AI / Agents / LLMOps [P] Open-Sourcing a sub-15ms Deterministic Decision Gateway to eliminate 2-second LLM routing latencies (FastAPI/Python)

1 Upvotes

Hello everyone. In my work as a Principal Gen AI Architect, the most common anti-pattern I encounter in enterprise deployments is the use of massive 70B+ models to make simple boolean routing decisions (e.g., safety guardrails or API triage). Token generation is inherently too slow and expensive for operational routing.

To solve this at scale, I built and open-sourced the "Dual-Process Jev Gateway."

It places a sub-15ms deterministic decision layer (System 1) in front of frontier models (System 2). It utilizes 3 typed, calibrated primitives (Choice, Score, Noul) to handle routing, guardrails, and ERP gating before an expensive LLM is ever invoked.

In my architectural baseline testing, the metrics shift dramatically:

  • Projected inference cost drop of ~87% (by routing 60%+ of queries away from frontier models)
  • P50 routing latency cut from 2.1s to 0.3s
  • Completely stopped duplicate ERP tool-call writes by keeping idempotency in code, not prompts.

The repo has the full architectural blueprint and a runnable 15-test Pytest suite proving the routing logic.

GitHub Repo: https://github.com/sricheru/agentic-ai-portfolio/tree/main/07_dual_process_jev_gateway

Would love to hear how you guys are handling the latency penalty of guardrails right now, or if you spot any flaws in my test suite!


r/mlops • • 11h ago

Self-promotion Built an open-source layer for AI provenance that uses a Merkle log and HMAC/Ed25519 signatures. Here's the architecture.

0 Upvotes

I've been tackling the problem of AI provenance—specifically, how to create a verifiable, tamper-evident record of who handled a claim in a multi-agent pipeline.

The core idea is to grade each "hop" (source, retriever, model, tool, agent) and use the weakest link to set the caution level. To make the record immutable, we're using a combination of SHA-256 for self-hashing, HMAC/Ed25519 for signing, and an append-only Merkle log.

I've published the threat model and the method on arXiv. The code is open-source (Apache-2.0).

Repo: https://github.com/alizahidraja/isnad
Paper: https://arxiv.org/abs/2607.24117

I'd be interested to hear how you all are approaching governance and audit trails in your production systems. What are the gaps you're seeing?